Showing 30 open source projects for "speak"

View related business solutions
  • One Monitoring Tool for IT, OT and Cloud | Free Trial Icon
    One Monitoring Tool for IT, OT and Cloud | Free Trial

    Vendor-agnostic monitoring across on-prem servers, cloud platforms and OT devices, all in one dashboard. No more tool sprawl.

    Modern infrastructure spans data centers, cloud platforms and factory floors, and every blind spot between them is a risk. PRTG supports SNMP, WMI, SSH and other standard protocols to monitor IT, OT and hybrid environments through one customizable dashboard. Build the views your team needs, from network health to application performance, without switching tools. Try PRTG free for 30 days now.
    Try PRTG Free
  • Build Securely on Azure with Proven Frameworks Icon
    Build Securely on Azure with Proven Frameworks

    Lay a foundation for success with Tested Reference Architectures developed by Fortinet’s experts. Learn more in this white paper.

    Moving to the cloud brings new challenges. How can you manage a larger attack surface while ensuring great network performance? Turn to Fortinet’s Tested Reference Architectures, blueprints for designing and securing cloud environments built by cybersecurity experts. Learn more and explore use cases in this white paper.
    Download Now
  • 1
    pyttsx3

    pyttsx3

    Offline Text To Speech synthesis for python

    ...The library exposes a simple but flexible API for controlling voice selection, speaking rate, volume, and other synthesis parameters from Python code. It supports both a high-level speak convenience function and a lower-level engine object with event hooks, queuing, and saving output to audio files. The repository includes examples and documentation that show how to adjust properties dynamically, persist synthesized output, and integrate pyttsx3 into GUIs or background services.
    Downloads: 20 This Week
    Last Update:
    See Project
  • 2
    JoyAI-VL-Interaction

    JoyAI-VL-Interaction

    An Open Real-time Video-Language Interaction System

    JoyAI-VL-Interaction is an open real-time video-language interaction system built around an 8B-scale vision-first model. It is designed to watch a webcam or livestream continuously and decide whether to speak, stay silent, or delegate a harder task. Unlike turn-based assistants, it focuses on event-driven interaction where timing matters as much as answer quality. The repository releases the model, training recipe, time-aligned interaction data, and deployable system together. Its system includes inference, WebUI, ASR, TTS, and background-agent services running on vLLM-based infrastructure. ...
    Downloads: 2 This Week
    Last Update:
    See Project
  • 3
    Open-LLM-VTuber

    Open-LLM-VTuber

    Open source AI VTuber platform with voice chat and Live2D avatars

    ...It enables hands-free conversations with large language models by combining speech recognition, language processing, and text-to-speech synthesis into a single system. Users can speak directly to the AI character, and the system can respond with a generated voice while animating a Live2D avatar to simulate a talking virtual personality. Open-LLM-VTuber is modular, allowing developers to swap or configure different language models, speech recognition engines, and voice synthesis systems depending on their needs. ...
    Downloads: 17 This Week
    Last Update:
    See Project
  • 4
    Real-Time AI Voice Chat

    Real-Time AI Voice Chat

    Have a natural, spoken conversation with AI

    ...RealtimeTTS then synthesizes the answer and streams speech back to the browser. Dynamic silence detection improves turn taking, while interruption handling lets users speak over an ongoing response. Multiple speech engines, including Kokoro, Coqui, and Orpheus, can be selected. Docker Compose simplifies deployment, although the original maintainer now treats the project as community-driven rather than actively developed.
    Downloads: 1 This Week
    Last Update:
    See Project
  • Ship Agents Faster Icon
    Ship Agents Faster

    Transform your applications and workflows into powerful agentic systems at global scale.

    Gemini Enterprise Agent Platform lets you rapidly build, scale, govern and optimize production-ready agents grounded in your organization's data. The platform enables developers to build custom or pre-built agents for virtually any use case. New customers get $300 in free credits.
    Start Free
  • 5
    OpenVoice

    OpenVoice

    Instant voice cloning by MIT and MyShell. Audio foundation model

    ...It is designed not only to match the timbre of the reference voice, but also to give granular control over style parameters such as emotion, accent, rhythm, pauses, and intonation. The model supports cross-lingual and even zero-shot cross-lingual voice cloning, so a speaker recorded in one language can be made to speak naturally in others. Architecturally, OpenVoice separates “tone color” cloning from style control, which makes it easier to keep a consistent identity while flexibly changing prosody or language. The project provides open-weight models, inference code, and examples, making it suitable both for research and for building production voice experiences. ...
    Downloads: 11 This Week
    Last Update:
    See Project
  • 6
    Fun Audio Chat

    Fun Audio Chat

    Large Audio Language Model built for natural interactions

    Fun Audio Chat is an interactive voice-first conversational AI platform designed to let users engage in natural spoken dialogue with large language models in real time, turning speech into context-aware responses while maintaining a smooth back-and-forth experience. It combines speech recognition, audio processing, and AI generation so users can speak simply and receive spoken replies, enabling applications such as virtual assistants, voice bots, and hands-free chat interfaces. The system supports dynamic audio input and output, meaning it can handle different voices, tones, and conversational contexts without forcing users into typed interactions. With real-time streaming, it minimizes latency and delivers responses quickly, making it suitable for applications where responsiveness matters, such as interactive demos, accessibility tools, and conversational games.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 7
    ChatterBot

    ChatterBot

    Machine learning, conversational dialog engine for creating chat bots

    ...For more details about the ideas and concepts behind ChatterBot see the process flow diagram. The language independent design of ChatterBot allows it to be trained to speak any language. Additionally, the machine-learning nature of ChatterBot allows an agent instance to improve it’s own knowledge of possible responses as it interacts with humans and other sources of informative data. An untrained instance of ChatterBot starts off with no knowledge of how to communicate. Each time a user enters a statement, the library saves the text that they entered and the text that the statement was in response to. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 8
    Step-Audio-EditX

    Step-Audio-EditX

    LLM-based Reinforcement Learning audio edit model

    ...Because the model is trained with a “large-margin learning” objective over many synthesized and natural speech samples, it gains robust control over expressive attributes, and can perform iterative editing: e.g. you could record a line, then ask the model to “make it sadder,” “speak slower,” or “change accent to X.”
    Downloads: 0 This Week
    Last Update:
    See Project
  • 9
    SpeakLogPSU
    SpeakLogPSU can speak chat messages with an individual voice if the NPC or player was configured or with a default one. You will never miss if someone talks to you. Voice cloning can be accomplished with Coqui in less than five minutes without GPU. The result is archived and can be used the next time in game. Some TTS projects already started to add tag support to speak text with emotions or sing it.
    Downloads: 0 This Week
    Last Update:
    See Project
  • Train ML Models With SQL You Already Know Icon
    Train ML Models With SQL You Already Know

    BigQuery automates data prep, analysis, and predictions with built-in AI assistance.

    Build and deploy ML models using familiar SQL. Automate data prep with built-in Gemini. Query 1 TB and store 10 GB free monthly.
    Start Free
  • 10
    translate-gui

    translate-gui

    GUI for translate-shell, the cli tool for quick translation

    GUI for the translate-shell, aims to be easy to use translator and a helpful tool for learning new languages. Most tools do a one way translation from source to target language, do to the reverse involves choosing the source and target languages again. This tool can do a 2 way translation accompanied by speech output of the target language text. Hence it can prove to be an indispensable aid when learning new languages
    Downloads: 1 This Week
    Last Update:
    See Project
  • 11
    Seadanse for Blender

    Seadanse for Blender

    Clay pass to Seedance 2.5 video, keeping the camera you keyframed

    ...The extension installs from disk into Blender 4.2 and up and lives in the 3D view's N-panel. The Agent Skill extracts into an AI assistant's skills directory and works with Claude Code, Codex, or anything that can run a shell and speak MCP. The server prices every job before a credit is spent, and the number sits on the Generate button. Appearance images control the look and add no credits. Requires a Seadanse account at https://seadanse.com. The reference guides the model; it does not guarantee exact reconstruction.
    Downloads: 1 This Week
    Last Update:
    See Project
  • 12
    PyNotes

    PyNotes

    An advanced Emacs-like text editor and IDE made in Python.

    PyNotes is an advanced Emacs-like text editor and IDE made in Python, but much simpler for new users.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 13
    EmotiVoice

    EmotiVoice

    Multi-Voice and Prompt-Controlled TTS Engine

    ...It supports both English and Chinese and ships with over 2,000 preset voices, making it suitable for everything from characters and virtual anchors to narration and dialogue. The core idea is prompt-based emotional and style control: you can ask the engine to speak “happy,” “sad,” “excited,” or with other high-level style prompts that shape prosody, pitch, speed, and energy. EmotiVoice provides multiple ways to interact with it, including a web interface, a Docker image, an HTTP API (including an OpenAI-compatible TTS API), and Python scripts for batch synthesis. It also supports voice cloning with your own data, backed by recipes for popular datasets like DataBaker and LJSpeech, so you can train or adapt voices to custom personas.
    Downloads: 1 This Week
    Last Update:
    See Project
  • 14
    VALL-E X

    VALL-E X

    Open source implementation of Microsoft's VALL-E X zero-shot TTS model

    ...The model attempts to match not just timbre, but also tone, pitch, emotion, and prosody of the reference audio, resulting in highly personalized output. VALL-E-X supports zero-shot cross-lingual synthesis, meaning a monolingual speaker’s voice can be used to speak other languages without additional training. It also preserves aspects of the acoustic environment, such as background noise or reverb, making the generated audio feel more like it came from the same setting as the prompt. The repository includes Python APIs, sample scripts, ready-to-use voice presets, and demos hosted on Hugging Face Spaces and Google Colab so users can try it.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 15
    alfred-ai

    alfred-ai

    The development of my ai assistant, Alfred

    The development of my ai assistant, Alfred.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 16
    ConvertText

    ConvertText

    ConvertText lets anyone change the case of your text easily.

    ConvertText is a very handy text tool where you can change between cases like Uppercase, Title Case, Random Case, Reverse Text, Backwards Text, Phonetic Alphabet and 1337 (Leet Speak).
    Downloads: 1 This Week
    Last Update:
    See Project
  • 17
    SVoice (Speech Voice Separation)

    SVoice (Speech Voice Separation)

    We provide a PyTorch implementation of the paper Voice Separation

    SVoice is a PyTorch-based implementation of Facebook Research’s study on speaker voice separation as described in the paper “Voice Separation with an Unknown Number of Multiple Speakers.” This project presents a deep learning framework capable of separating mixed audio sequences where several people speak simultaneously, without prior knowledge of how many speakers are present. The model employs gated neural networks with recurrent processing blocks that disentangle voices over multiple computational steps, while maintaining speaker consistency across output channels. Separate models are trained for different speaker counts, and the largest-capacity model dynamically determines the actual number of speakers in a mixture. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 18
    Multilingual Speech Synthesis

    Multilingual Speech Synthesis

    An implementation of Tacotron 2 that supports multilingual experiments

    ...It contains an implementation of Tacotron 2 that supports multilingual experiments and that implements different approaches to encoder parameter sharing. It presents a model combining ideas from Learning to speak fluently in a foreign language: Multilingual speech synthesis and cross-language voice cloning, End-to-End Code-Switched TTS with Mix of Monolingual Recordings, and Contextual Parameter Generation for Universal Neural Machine Translation. We provide data for comparison of three multilingual text-to-speech models. The first shares the whole encoder and uses an adversarial classifier to remove speaker-dependent information from the encoder. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 19

    SmartBody

    Character animation system for games and simulations.

    ...SmartBody is a character animation platform that provides the following capabilities in real time: * Locomotion (walk, jog, run, turn, strafe, jump, etc.) * Steering - avoiding obstacles and moving objects * Object manipulation - reach, grasp, touch , pick up objects * Lip Syncing - characters can speak with simultaneous lip-sync using text-to-speech or prerecorded audio * Gazing - robust gazing behavior that incorporates various parts of the body * Nonverbal behavior - gesturing, head nodding and shaking, eye saccades - Online and offline retargeting of motion - Automatic skinning and rigging SmartBody is written in C++ and can be incorporated into most game engines. ...
    Downloads: 6 This Week
    Last Update:
    See Project
  • 20
    Mollino

    Mollino

    Not your usual Architectural Modeler

    ...As I am an architect, and I know very little about programming, and wouldn't reach even in 10 years the necessary level to be able to write anything useful for this type of software, my part will be to bring ideas and coherence to this project. If you want to know some more : - go to the blog https://sourceforge.net/p/mollino/blog/ - write me a message ;) I speak french, german, english and italian. Martin Lucas
    Downloads: 0 This Week
    Last Update:
    See Project
  • 21
    Voice keyboard/dictation. Aims to be a total substitute for a keyboard. Spell out words letter by letter (using code: alpha, bravo, ..). Arrow keys, modifiers work. Speak whole words (but whole word accuracy is not good). Attach commands to some word
    Downloads: 0 This Week
    Last Update:
    See Project
  • 22
    feinerleiser

    feinerleiser

    Feinerleiser implements cite rules from the humanities for jabref

    Feinerleiser (speak finalizer) takes an odt file that has citations inserted by jabref and expands them according to rules set by the user in an xml file. Contrary to the style of citation used in engineering, psychology or medicine [Doe 1985; 123] the user can create complex citation-styles, that satisfy the needs and expectations of researchers from the field of humanities.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 23
    Geotran is a tool meant to help geocachers who are traveling to a foreign country, but do not speak the language. Geotran, is a python tool that invokes Google Translator to translate the pocket queries.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 24
    This is my first source project please do not rip me off and steal my project, Its Open Source now in English and Spanish, I speak English and english only so spanish is translated so i might now make another spanish one again.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 25
    A lightweight python twitter reader with speech and growl support, designed to run in the background and let you get tweets in realtime, without having to run off to the site or watch an update feed constantly.
    Downloads: 0 This Week
    Last Update:
    See Project
  • Previous
  • You're on page 1
  • 2
  • Next