Showing 85 open source projects for "speech engine"

View related business solutions
  • Demo Series - Small Business Backup By Veeam Icon
    Demo Series - Small Business Backup By Veeam

    Learn how to protect your Microsoft 365 data, with simple, actionable tips today.

    Watch this on-demand demo series and learn how to protect your Microsoft 365 data with clear, simple, actionable steps that are easy to implement for businesses of all sizes.
    Watch Demo Series
  • $300 Free Credits to Build on Google Cloud Icon
    $300 Free Credits to Build on Google Cloud

    New customers can spin up VMs, build with AI, and query data at no cost.

    Put your $300 in credit toward real workloads, then keep building with free monthly usage for 20+ products. No commitment and no charge until you upgrade.
    Start Free
  • 1
    kokoro-onnx

    kokoro-onnx

    TTS with kokoro and onnx runtime

    kokoro-onnx is a text-to-speech toolkit that wraps the Kokoro neural TTS model in an easy-to-use ONNX Runtime interface, so you can generate speech from Python with minimal setup. It focuses on running efficiently on commodity hardware, including macOS with Apple Silicon, while still delivering near real-time performance for many use cases. The project ships prebuilt model files and a simple example script, so you can go from installation to producing an audio.wav file in just a few steps....
    Downloads: 316 This Week
    Last Update:
    See Project
  • 2
    pyttsx3

    pyttsx3

    Offline Text To Speech synthesis for python

    pyttsx3 is an offline text-to-speech library for Python that wraps native speech engines instead of calling cloud APIs. It is designed to work entirely without an internet connection, making it suitable for local automation, kiosks, accessibility tools, and embedded applications. On Windows it uses SAPI5, on Linux it typically uses eSpeak or eSpeak-NG, and on macOS it can use NSSpeechSynthesizer or AVSpeechSynthesizer, giving it broad cross-platform compatibility.
    Downloads: 23 This Week
    Last Update:
    See Project
  • 3
    ESPnet

    ESPnet

    End-to-end speech processing toolkit

    ESPnet is a comprehensive end-to-end speech processing toolkit covering a wide spectrum of tasks, including automatic speech recognition (ASR), text-to-speech (TTS), speech translation (ST), speech enhancement, speaker diarization, and spoken language understanding. It uses PyTorch as its deep learning engine and adopts a Kaldi-style data processing pipeline for features, data formats, and experimental recipes.
    Downloads: 7 This Week
    Last Update:
    See Project
  • 4
    abogen

    abogen

    Generate audiobooks from EPUBs, PDFs and text with captions

    abogen is a tool designed to generate audiobooks (or speech narrations) from textual sources such as EPUBs, PDFs, or plain text, with synchronized captions. In other words, it automates the pipeline of reading a digital book (or document), converting its text into speech via a TTS engine, and packaging the result into an audiobook format — likely along with timestamped captions or subtitles that align with the spoken audio.
    Downloads: 10 This Week
    Last Update:
    See Project
  • MongoDB Atlas runs apps anywhere Icon
    MongoDB Atlas runs apps anywhere

    Deploy in 115+ regions with the modern database for every enterprise.

    MongoDB Atlas gives you the freedom to build and run modern applications anywhere—across AWS, Azure, and Google Cloud. With global availability in over 115 regions, Atlas lets you deploy close to your users, meet compliance needs, and scale with confidence across any geography.
    Start Free
  • 5
    RealtimeSTT

    RealtimeSTT

    A robust, efficient, low-latency speech-to-text library

    RealtimeSTT is a Python-based realtime speech-to-text engine emphasizing low latency, wake-word detection, voice activity detection, and automatic speech segmentation. It provides asynchronous callbacks, nanosecond-precision timestamps, and CLI tools, suitable for building voice assistants, meeting transcribers, or live caption systems.
    Downloads: 1 This Week
    Last Update:
    See Project
  • 6
    Voice-Pro

    Voice-Pro

    Comprehensive Gradio WebUI for audio processing

    Voice-Pro is the best gradio WebUI for transcription, translation and text-to-speech. It can be easily installed with one click. Create a virtual environment using Miniconda, running completely separate from the Windows system (fully portable). Supports real-time transcription and translation, as well as batch mode.
    Downloads: 16 This Week
    Last Update:
    See Project
  • 7
    RealtimeTTS

    RealtimeTTS

    Converts text to speech in realtime

    RealtimeTTS is a low-latency text-to-speech library built for real-time applications such as voice chat with LLMs, assistants, and interactive tools. It is designed around a streaming model: you can feed it text incrementally (for example, as an LLM responds) and get audio output almost immediately, which keeps end-to-end latency very low. The library is engine-agnostic and plugs into a wide range of cloud and local TTS systems, including OpenAI, ElevenLabs, Azure, Coqui, Piper, StyleTTS2, Edge TTS, Google TTS, system TTS and others, so you can swap providers without rewriting your pipeline. ...
    Downloads: 8 This Week
    Last Update:
    See Project
  • 8
    AI Runner

    AI Runner

    Offline inference engine for art, real-time voice conversations

    AI Runner is an offline inference engine designed to run a collection of AI workloads on your own machine, including image generation for art, real-time voice conversations, LLM-powered chatbots and automated workflows. It is implemented as a desktop-oriented Python application and emphasizes privacy and self-hosting, allowing users to work with text-to-speech, speech-to-text, text-to-image and multimodal models without sending data to external services.
    Downloads: 1 This Week
    Last Update:
    See Project
  • 9
    Qwen3-ASR

    Qwen3-ASR

    Qwen3-ASR is an open-source series of ASR models

    Qwen3-ASR is an automatic speech recognition system in the QwenLM family, developed to convert spoken language into text with strong accuracy and real-time performance. As a specialized ASR variant of the broader Qwen language model ecosystem, it focuses on capturing reliable transcriptions from audio sources such as recordings, live streams, or conversational inputs while supporting low latency use cases. The architecture combines advanced neural acoustic modeling with context-aware...
    Downloads: 2 This Week
    Last Update:
    See Project
  • Paessler - Monitor Your Whole Network in Minutes Icon
    Paessler - Monitor Your Whole Network in Minutes

    Auto-discovery finds your devices and deploys pre-configured sensors instantly. No project plan required, just visibility from day one.

    Waiting weeks for a monitoring rollout isn't an option when infrastructure doesn't stop running. PRTG's auto-discovery scans your network and suggests from over 200 pre-configured sensor types, so you're watching servers, applications and devices within minutes, not after a multi-week deployment. Enterprise-strength monitoring, without the enterprise complexity. Start your free trial today.
    Start Free 30-Day Trial
  • 10
    EasyVoice

    EasyVoice

    Open source text-to-speech tool, supports extra-long text

    easyVoice is an open-source text-to-speech platform aimed at turning long-form text and novels into high-quality audio, with a strong focus on usability and scalability. It provides a web interface where users can paste or upload large texts and generate speech and subtitles in a single workflow, even for works exceeding 100,000 characters. The system supports multi-role voice acting, letting users assign different neural voices to different characters or narrative roles and configure...
    Downloads: 4 This Week
    Last Update:
    See Project
  • 11
    Concat

    Concat

    Open-Source CapCut replacement (MCP supported)

    Concat is a free and open-source desktop video editor positioned as a local alternative to subscription-based editors such as CapCut. It uses a native Rust engine and performs editing entirely on the user’s machine. Projects can contain multiple tracks and timelines for more complex compositions. Core editing tools include splitting, trimming, merging, transitions, and playback-speed adjustment. Offline features include automatic captions, local text-to-speech voices, voice processing, titles, and styled text. ...
    Downloads: 65 This Week
    Last Update:
    See Project
  • 12
    Transcoder

    Transcoder

    Hardware-accelerated video transcoding using Android MediaCodec APIs

    Transcoder by DeepMedia is an AI-powered video-to-video speech translation engine that enables fully automated multilingual dubbing. Unlike traditional speech translation systems that rely on multi-stage pipelines, Transcoder directly translates one speaker’s video into another language while preserving facial expressions, lip-sync, and vocal identity. Designed for real-time use and production-grade pipelines, Transcoder combines advanced deep learning models with GPU acceleration to deliver high-quality translations across languages. ...
    Downloads: 8 This Week
    Last Update:
    See Project
  • 13
    Smile

    Smile

    Statistical machine intelligence and learning engine

    Smile is a fast and comprehensive machine learning engine. With advanced data structures and algorithms, Smile delivers the state-of-art performance. Compared to this third-party benchmark, Smile outperforms R, Python, Spark, H2O, xgboost significantly. Smile is a couple of times faster than the closest competitor. The memory usage is also very efficient. If we can train advanced machine learning models on a PC, why buy a cluster? Write applications quickly in Java, Scala, or any JVM...
    Downloads: 21 This Week
    Last Update:
    See Project
  • 14
    Rhino

    Rhino

    On-device Speech-to-Intent engine powered by deep learning

    Rhino is Picovoice's Speech-to-Intent engine. It directly infers intent from spoken commands within a given context of interest, in real-time. The end-to-end platform for embedding private voice AI into any software in a few lines of code. Design with no limits on top of a modular platform. Create use-case-specific voice AI models in seconds. Develop voice features with a few lines of code using intuitive and cross-platform SDKs.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 15
    gse

    gse

    Go efficient multilingual NLP and text segmentation

    Go efficient multilingual NLP and text segmentation; support English, Chinese, Japanese and others. Gse is implements jieba by golang, and try add NLP support and more feature. Support common, search engine, full mode, precise mode and HMM mode multiple word segmentation modes. Support user and embed dictionary, Part-of-speech/POS tagging, analyze segment info, stop and trim words. Support multilingual: English, Chinese, Japanese and others. Support Traditional Chinese. Support HMM cut text use Viterbi algorithm. Support NLP by TensorFlow (in work). ...
    Downloads: 1 This Week
    Last Update:
    See Project
  • 16
    Simple TTS Reader

    Simple TTS Reader

    A small clipboard reader

    Simple TTS Reader is a small utility that reads text from your clipboard using Microsoft Speech API. Whenever you copy any text, the app instantly converts it into spoken words. Select your preferred speech engine from those installed on your system, such as Microsoft Zira, and adjust speed and volume for personalized playback. The application can also be minimized to the system tray. Plus, it is free and comes with an intuitive interface that makes it accessible to everyone.
    Leader badge
    Downloads: 56 This Week
    Last Update:
    See Project
  • 17
    Scribe

    Scribe

    Free, open-source, and offline speech-to-text & voice control app.

    ...Powered by the Vosk engine, it supports multiple languages and provides high-quality recognition without an internet connection. > Scribe is the perfect tool for anyone looking to boost productivity, improve accessibility, or simply interact with their computer in a new, hands-free way.
    Downloads: 82 This Week
    Last Update:
    See Project
  • 18
    JSpeech

    JSpeech

    Java library designed to integrate Speech-to-Text

    jSpeech is a Java library designed to integrate Speech-to-Text (STT) capabilities, command control, and diarization (speaker identification) into applications in a simple, modular, and decoupled way.
    Downloads: 6 This Week
    Last Update:
    See Project
  • 19
    Universal-TTS-Pro

    Universal-TTS-Pro

    Portable, fully offline Text-to-Speech (TTS) supporting 30+ languages

    Fast Download: https://github.com/szabiz/Universal-TTS-Pro/releases/download/1.3.4/UniversalTTS_Pro_1.3.4.zip • Universal TTS Pro: Converts text to speech offline. • Dual TTS Engine: Piper & Supertonic (100% private, local processing). • Smart Normalization & Number-to-Words supporting Hungarian (HU), English (EN), and Romanian (RO). • Max Chunk Length control, Volume settings, and Supertonic 3 phonetic corrections. • Clean, smaller executable build with improved audio export quality (no clipping)...
    Downloads: 3 This Week
    Last Update:
    See Project
  • 20
    AudioBC

    AudioBC

    Offline desktop app to convert EPUB to MP3 using Kokoro-82M neural TTS

    AudioBC is a powerful desktop application designed to turn your digital library into a personal audiobook collection. Unlike most Text-to-Speech (TTS) tools that require expensive cloud API subscriptions or an active internet connection, AudioBC runs entirely on your local machine. Powered by the state-of-the-art Kokoro-82M neural engine, AudioBC produces natural, human-like speech that rivals premium cloud services. It is built with a focus on privacy and simplicity, offering a streamlined three-step workflow: Extract, Edit, and Convert. ...
    Downloads: 1 This Week
    Last Update:
    See Project
  • 21
    Eva AI

    Eva AI

    Eva is an A.I. assistant that helps users multi-task.

    ...For more instructions, check the instruction manual included in the application. [Update] * 🆕 Eva stays activated on 'Listen' and stops on 'Stop Listening' * 🆕 Added a new on-device speech recognition engine: Moonshine 🌕 * 🆕 Added a new on-device speech synthesis engine: Piper 🪈 * 🆕 Removed recognition timeouts. Commands can now have an indefinite duration. * 🆕 Removed unnecessary STT settings * 🐞 Fixed command customisation bugs
    Downloads: 7 This Week
    Last Update:
    See Project
  • 22
    pdf-to-podcast

    pdf-to-podcast

    PDF to Podcast transforms any PDF document into a podcast-ready audio

    PDF to Podcast transforms any PDF document into a podcast-ready audio episode using advanced AI text-to-speech (TTS) providers. Upload a PDF, select your preferred voice and provider, and receive an MP3 and a ready-to-use RSS feed for your podcast app.
    Downloads: 1 This Week
    Last Update:
    See Project
  • 23
    PearlOS

    PearlOS

    Latest Debian and Ubuntu based REPO files and ISO's for PearlLinuxOS

    PearlOS is Pearl Linux just another branding we used early on with the distro. We are hoping to better organize our upcoming releases. All future releases and the most resent scootski (12) and Preslee (13) you will find under pearlos. Our previous release of our first rolling release also called preslee is no longer on bookworm and trixie now its stickly trixie which is our latest debian release. Sccotski is our latest Ubuntu base and is from 24.04 ISO's are located under...
    Leader badge
    Downloads: 209 This Week
    Last Update:
    See Project
  • 24
    ShortGeek

    ShortGeek

    Free short video maker for Windows - article in, narrated short out

    ShortGeek turns one of your guides, an RSS feed or a bare topic into a narrated, captioned vertical short, rendered on your own machine. Pick a source, review the script before anything renders, choose a voice and a look, then let it render. Hook, beats and call to action are all editable first. Captions are word level and timed from the voice engine rather than estimated, in three styles. Eight backgrounds are generated in code, frame by frame, so nothing is AI art, and you can drop...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 25
    TamilDictation-Plugin

    TamilDictation-Plugin

    A simple Microsoft Word add-in to type, translate, and format Tamil te

    Tamil Dictation is a Microsoft Word Task Pane Add-in that enables native speakers, government officials, and content creators to dictate directly in Tamil or English — right inside Word. It uses the browser's Web Speech API for voice recognition and the Office.js API to write text precisely at the cursor position, with no copy-paste needed. The add-in runs locally via a development server and is sideloaded into Word using a manifest.xml file — meaning there is no standalone application to...
    Downloads: 0 This Week
    Last Update:
    See Project
  • Previous
  • You're on page 1
  • 2
  • 3
  • 4
  • Next