Large Language Model Text Generation Inference
Instant voice cloning by MIT and MyShell. Audio foundation model
Spark-TTS Inference Code
Multi-modal large language model designed for audio understanding
Googles NotebookLM but local
First class Sublime Text AI assistant with gpt-5, Opus 4.6, Gemini 3
End-to-end speech processing toolkit
A PyTorch-based Speech Toolkit
AI tool that removes hardcoded subtitles and text from videos locally
A GUI tool for extracting hard-coded subtitle (hardsub) from videos
TTS with kokoro and onnx runtime
Python tool for converting files and office documents to Markdown
SOTA Open Source TTS
OCRmyPDF adds an OCR text layer to scanned PDF files
A text-to-speech, speech-to-text and speech-to-speech library
MiniMax H3 is a general-purpose, omni-modal generative system
Open-Source Python3 tool for recognizing layouts, tables, and math
Awesome multilingual OCR toolkits based on PaddlePaddle
Jupyter Notebooks as Markdown Documents, Julia, Python or R scripts
Comprehensive Gradio WebUI for audio processing
Qwen3-TTS is an open-source series of TTS models
A TTS that fits in your CPU (and pocket)
Strip multi-vendor AI provenance marks
Wan2.2: Open and Advanced Large-Scale Video Generative Model
Wan2.1: Open and Advanced Large-Scale Video Generative Model