A GUI tool for extracting hard-coded subtitle (hardsub) from videos
A robust, efficient, low-latency speech-to-text library
Official inference repo for FLUX.1 models
Comprehensive Gradio WebUI for audio processing
Generate audiobooks from EPUBs, PDFs and text with captions
State-of-the-art TTS model under 25MB
A TTS that fits in your CPU (and pocket)
A simple native web interface that uses ChatTTS to synthesize text
EPUB to audiobook converter, optimized for Audiobookshelf
The behavior guidance framework for customer-facing LLM agents
Offline inference engine for art, real-time voice conversations
Automatic Speech Recognition with Word-level Timestamps
A nearly-live implementation of OpenAI's Whisper
Robust Speech Recognition via Large-Scale Weak Supervision
Code for running inference and finetuning with SAM 3 model
A lightweight text-to-speech model with zero-shot voice cloning
OCR software, free and offline
Contexts Optical Compression
Qwen3-TTS is an open-source series of TTS models
An open-source toolkit for monitoring Language Learning Models (LLMs)
Speech recognition module for Python
Video-based AI memory library. Store millions of text chunks in MP4
SOTA Open Source TTS
State-of-the-art Machine Learning for Pytorch, TensorFlow, and JAX
Google Gen AI Python SDK provides an interface for developers