Showing 353 open source projects for "speech"

View related business solutions
  • $300 Free Credits to Build on Google Cloud Icon
    $300 Free Credits to Build on Google Cloud

    New customers can spin up VMs, build with AI, and query data at no cost.

    Put your $300 in credit toward real workloads, then keep building with free monthly usage for 20+ products. No commitment and no charge until you upgrade.
    Start Free
  • One Monitoring Tool for IT, OT and Cloud | Free Trial Icon
    One Monitoring Tool for IT, OT and Cloud | Free Trial

    Vendor-agnostic monitoring across on-prem servers, cloud platforms and OT devices, all in one dashboard. No more tool sprawl.

    Modern infrastructure spans data centers, cloud platforms and factory floors, and every blind spot between them is a risk. PRTG supports SNMP, WMI, SSH and other standard protocols to monitor IT, OT and hybrid environments through one customizable dashboard. Build the views your team needs, from network health to application performance, without switching tools. Try PRTG free for 30 days now.
    Try PRTG Free
  • 1
    MARS5

    MARS5

    MARS5 speech model (TTS) from CAMB.AI

    MARS5-TTS is CAMB.AI’s open-source English speech model designed for high-quality text-to-speech and voice emulation. It uses a two-stage architecture that combines an autoregressive (AR) model with a non-autoregressive (NAR) model, giving it both expressiveness and speed. The model is built to handle prosodically challenging content such as sports commentary, anime dialogue, and other high-energy or highly varied speech patterns with realistic rhythm and intonation. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 2
    KoboldCpp

    KoboldCpp

    Run GGUF models easily with a UI or API. One File. Zero Install.

    KoboldCpp is an easy-to-use AI text-generation software for GGML and GGUF models, inspired by the original KoboldAI. It's a single self-contained distributable that builds off llama.cpp and adds many additional powerful features.
    Leader badge
    Downloads: 347 This Week
    Last Update:
    See Project
  • 3
    Auto-CS

    Auto-CS

    Automatic annotation of Cued Speech

    Auto-CS is a Python program developed within the project AutoCuedSpeech: https://auto-cuedspeech.org. Auto-CS contains all the components dedicated to the automatic annotation of Cued Speech. This source code is not a standalone tool: it runs exclusively within SPPAS. It is integrated through the **spin-off** mechanism provided by SPPAS, which allows external code bases to remain separate while still being dynamically discovered and used by the main framework. See <https://auto-cuedspeech.org> for details about the related Research Project. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 4

    WhisperJAV

    A subtitle generator for Japanese Adult Videos.

    A subtitle generator for Japanese Adult Videos. Transformer-based ASR architectures like Whisper suffer significant performance degradation when applied to the spontaneous and noisy domain of JAV. This degradation is driven by specific acoustic and temporal characteristics that defy the statistical distributions of standard training data.
    Downloads: 61 This Week
    Last Update:
    See Project
  • 5
    Eva AI

    Eva AI

    Eva is an A.I. assistant that helps users multi-task.

    ...For more instructions, check the instruction manual included in the application. [Update] * 🆕 Eva stays activated on 'Listen' and stops on 'Stop Listening' * 🆕 Added a new on-device speech recognition engine: Moonshine 🌕 * 🆕 Added a new on-device speech synthesis engine: Piper 🪈 * 🆕 Removed recognition timeouts. Commands can now have an indefinite duration. * 🆕 Removed unnecessary STT settings * 🐞 Fixed command customisation bugs
    Downloads: 5 This Week
    Last Update:
    See Project
  • 6
    TITTSE

    TITTSE

    Two Integrated Text To Speech Engines uses MMS & Silero

    TITTSE is a Python Application that allows you to easily and quickly convert text to speech in 15 different languages (or add more easily) using Two TTS Engines. All you need is a text file ending in the tittse extension with 4 header lines including the TITTSE language code (see documentation for your language), the 'base' file name for the audio files TITTSE creates, voice gender (girl or boy), offset (file numbers added to base file name start at this number).
    Downloads: 0 This Week
    Last Update:
    See Project
  • 7
    translate-gui

    translate-gui

    GUI for translate-shell, the cli tool for quick translation

    ...Most tools do a one way translation from source to target language, do to the reverse involves choosing the source and target languages again. This tool can do a 2 way translation accompanied by speech output of the target language text. Hence it can prove to be an indispensable aid when learning new languages
    Downloads: 0 This Week
    Last Update:
    See Project
  • 8
    SpeakLogPSU
    SpeakLogPSU can speak chat messages with an individual voice if the NPC or player was configured or with a default one. You will never miss if someone talks to you. Voice cloning can be accomplished with Coqui in less than five minutes without GPU. The result is archived and can be used the next time in game. Some TTS projects already started to add tag support to speak text with emotions or sing it. If a game designer has that in mind with a good chat log she can voiced her...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 9
    Auditory Modeling Toolbox
    The auditory modeling toolbox (AMT) is a Matlab/Octave toolbox for the development and application of auditory computational models. Over 50 auditory models implemented in Matlab, Octave, C, C++, and Python can be run from Matlab and Octave, on Windows and Linux. The AMT provides a well-structured in-code documentation, includes auditory data required to run the models. It integrates functionality to reproduce the model predictions. Model implementations can be evaluated in two stages,...
    Leader badge
    Downloads: 31 This Week
    Last Update:
    See Project
  • MongoDB Atlas runs apps anywhere Icon
    MongoDB Atlas runs apps anywhere

    Deploy in 115+ regions with the modern database for every enterprise.

    MongoDB Atlas gives you the freedom to build and run modern applications anywhere—across AWS, Azure, and Google Cloud. With global availability in over 115 regions, Atlas lets you deploy close to your users, meet compliance needs, and scale with confidence across any geography.
    Start Free
  • 10
    AppGrabbit

    AppGrabbit

    Download, transcribe and convert videos on your PC. No cloud.

    AppGrabbit is a Windows 10/11 app that downloads, transcribes and converts videos on your own computer. Paste a link from YouTube (videos, Shorts, playlists, Mixes), TikTok, Instagram, Facebook, X, Pinterest, Twitch or SoundCloud and the file lands in your downloads folder in the quality you pick, up to 4K. Turn any video into a clean .txt with a prompt on top for ChatGPT, Claude or Gemini: with captions it is instant; without them, or from your own recording (a class, a meeting, a voice...
    Downloads: 3 This Week
    Last Update:
    See Project
  • 11
    ChatTTS_colab

    ChatTTS_colab

    One-click deployment (including offline integration package)

    ChatTTS_colab is a wrapper project around the ChatTTS model that focuses on “one-click” deployment, especially in Google Colab. It provides an integrated offline bundle and scripts for Windows and macOS so users can run ChatTTS locally without wrestling with complex environment setup. The repository includes Colab notebooks that launch a Gradio-based web UI and expose streaming TTS, making it possible to listen to generated audio as it is produced. A distinctive feature is the “voice gacha”...
    Downloads: 1 This Week
    Last Update:
    See Project
  • 12
    FranMsxApps

    FranMsxApps

    My MSX programs and some additional .cas tools

    ...origin=sourceforge Some of the programs are cool. The most relevant are: - A graphic designer in Assembler Z-80 - A ship game in Assembler Z-80 - A text to speech in Spanish. - The seed of a maze game. - My version of Tetris in Assembler Z-80 There are some demo videos, and even a video of a hisoft assembler session. I hope you enjoy
    Downloads: 0 This Week
    Last Update:
    See Project
  • 13
    ShortGeek

    ShortGeek

    Free short video maker for Windows - article in, narrated short out

    ShortGeek turns one of your guides, an RSS feed or a bare topic into a narrated, captioned vertical short, rendered on your own machine. Pick a source, review the script before anything renders, choose a voice and a look, then let it render. Hook, beats and call to action are all editable first. Captions are word level and timed from the voice engine rather than estimated, in three styles. Eight backgrounds are generated in code, frame by frame, so nothing is AI art, and you can drop...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 14
    anchorcastapp

    anchorcastapp

    Free AI-powered church presentation & live sermon display app

    AnchorCast is a free, open-source AI-powered church presentation desktop app for Windows and MacOS, built with Electron. Features: - Live Sermon Transcription — real-time speech-to-text via Whisper AI - AI Bible Verse Detection — automatically detects and displays verses from live sermons - Song Manager — display song lyrics on projection - Media Playback — images and video on projection screen - NDI Output — stream projection over local network - Remote Control — control presentation from any phone via Wi-Fi - Timer Display — countdown and service timers - Presentation Editor — create and manage slide decks - Theme Manager — customize projection appearance - Sermon Intelligence & Analytics Free for all churches. ...
    Downloads: 13 This Week
    Last Update:
    See Project
  • 15
    Cowser

    Cowser

    greets you with ASCII art animals. Inspired by cowsay

    Cowser is a command-line tool (and Python library) that wraps a message in a speech or thought bubble and prints it above a randomly chosen animal. It supports 35 built-in animals, cowsay-style moods, colors, fortunes, interactive REPL sessions, and user-supplied art packs — while staying a pure, side-effect-free function you can also call from your own code.
    Downloads: 1 This Week
    Last Update:
    See Project
  • 16
    Lingueez - Vocabulary Builder

    Lingueez - Vocabulary Builder

    Collect vocabulary from anywhere, organize, and master it.

    ...And when recognizing a word is no longer enough, the quiz asks you to pick the right answer — or type it out yourself, in either direction. Beyond flashcards, Lingueez reads your words aloud with natural text-to-speech and generates short reading passages built from the vocabulary you're learning, so you see every word in real context. Your library is yours: everything works offline, with no account required. Sign in only if you want your words and progress synced across devices. Free and open source. No ads, no subscription.
    Downloads: 15 This Week
    Last Update:
    See Project
  • 17
    PyNotes

    PyNotes

    An advanced Emacs-like text editor and IDE made in Python.

    PyNotes is an advanced Emacs-like text editor and IDE made in Python, but much simpler for new users.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 18

    Tokenized Text Aligner

    Aligns tokens in two versions of a text with differing tokenization.

    ...It is intended for use in the preparation of annotated linguistic corpora, where differences in tokenization may arise (i) following corrections or modifications to the source text or (ii) through the creation of different layers of annotation (part-of-speech, treebank) requiring different tokenization. In its default implementation, it produces a human-readable CSV table associating tokens in text A with tokens in text B, and can also inject token-level annotation from text B to text A. The Aligner class on which the default implementation is based can be incorporated into more complex workflows.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 19
    SoniTranslate

    SoniTranslate

    Synchronized Translation for Videos

    ...It provides a web UI built with Gradio, allowing users to upload a video, choose source and target languages, and then run a pipeline that handles transcription, translation and re-synthesis of speech. Under the hood, it uses advanced speech and diarization models to separate speakers, align audio with timecodes and respect subtitle timing, which lets the generated dub track stay in sync with the original video structure. The project supports a wide range of languages for translation, spanning major world languages (English, Spanish, French, German, Chinese, Arabic, etc.) and many regional or less widely spoken languages, making it suitable for broad internationalization. ...
    Downloads: 27 This Week
    Last Update:
    See Project
  • 20

    Mice TTM

    mice stt tts

    Dieses Tool wird speziell für die Barrierefreiheit unter Linux entwickelt. Es ermöglicht das umwandeln/konvertieren/parsen von Texten die aus einer Spracherkennung stammen, in Diktate sowie das Ausführen von Makros. Dies funktioniert ohne Internet, da die Spracherkennung auf dem PC selbst erfolgt. Mausbewegungen auf benannte Wörter und dann entsprechend auswählen oder per Sprachbefehl klicken. Außerdem können Textpassagen z.B. unter Libreoffice Wirter per Sprachbefehl entsprechend...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 21
    Insanely Fast Whisper

    Insanely Fast Whisper

    An opinionated CLI to transcribe Audio files w/ Whisper on-device

    Insanely Fast Whisper is a high-performance command-line tool designed to dramatically accelerate speech-to-text transcription using OpenAI’s Whisper models on local hardware. It leverages modern optimizations such as batch processing, mixed precision, and advanced attention mechanisms like Flash Attention to significantly reduce inference time while maintaining high transcription accuracy. The project is built on top of the Transformers ecosystem and integrates with libraries such as Optimum to maximize GPU efficiency. ...
    Downloads: 1 This Week
    Last Update:
    See Project
  • 22
    Vocode

    Vocode

    Build voice-based LLM agents. Modular + open source

    Vocode is an open source library that makes it easy to build voice-based LLM apps. Using Vocode, you can build real-time streaming conversations with LLMs and deploy them to phone calls, Zoom meetings, and more. You can also build personal assistants or apps like voice-based chess. Vocode provides easy abstractions and integrations so that everything you need is in a single library.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 23
    Luna AI

    Luna AI

    Virtual AI anchor that combines state-of-the-art technology

    ...The project supports multiple rendering backends for the avatar, such as Live2D, Unreal Engine (UE), and “xuniren,” and can output to streaming platforms like Bilibili, Douyin, Kuaishou, WeChat Channels, Pinduoduo, Douyu, YouTube, Twitch, and TikTok. For voice, it integrates with numerous TTS engines (Edge-TTS, VITS-Fast, ElevenLabs, VALL-E-X, OpenVoice, GPT-SoVITS, Azure TTS, fish-speech, ChatTTS, CosyVoice, F5-TTS, MultiTTS, MeloTTS, and others), and can optionally pass the output through voice conversion systems like so-vits-svc or DDSP-SVC to change timbre.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 24
    StyleTTS 2

    StyleTTS 2

    Towards Human-Level Text-to-Speech through Style Diffusion

    StyleTTS2 is a state-of-the-art text-to-speech system that aims for human-level naturalness by combining style diffusion, adversarial training, and large speech language models. It extends the original StyleTTS idea by introducing a style diffusion model that can sample rich, realistic speaking styles conditioned on reference speech, allowing highly expressive and diverse prosody.
    Downloads: 2 This Week
    Last Update:
    See Project
  • 25
    MeloTTS

    MeloTTS

    High-quality multi-lingual text-to-speech library by MyShell.ai

    MeloTTS is an open-source text-to-speech (TTS) system that generates natural-sounding speech from text input. It utilizes advanced machine-learning models to produce high-quality audio outputs.
    Downloads: 1 This Week
    Last Update:
    See Project