High-Quality Voice Cloning TTS for 600+ Languages
Awesome multilingual OCR toolkits based on PaddlePaddle
A GUI tool for extracting hard-coded subtitle (hardsub) from videos
A robust, efficient, low-latency speech-to-text library
Official inference repo for FLUX.1 models
Comprehensive Gradio WebUI for audio processing
Generate audiobooks from EPUBs, PDFs and text with captions
State-of-the-art TTS model under 25MB
A TTS that fits in your CPU (and pocket)
AI bridge enabling assistants to control and automate Unity Editor
A simple native web interface that uses ChatTTS to synthesize text
EPUB to audiobook converter, optimized for Audiobookshelf
Offline inference engine for art, real-time voice conversations
The behavior guidance framework for customer-facing LLM agents
Automatic Speech Recognition with Word-level Timestamps
A nearly-live implementation of OpenAI's Whisper
Code for running inference and finetuning with SAM 3 model
Robust Speech Recognition via Large-Scale Weak Supervision
A lightweight text-to-speech model with zero-shot voice cloning
OCR software, free and offline
Contexts Optical Compression
Qwen3-TTS is an open-source series of TTS models
An open-source toolkit for monitoring Language Learning Models (LLMs)
Speech recognition module for Python
Video-based AI memory library. Store millions of text chunks in MP4