OCR software, free and offline
A TTS that fits in your CPU (and pocket)
Contexts Optical Compression
SoTA open-source TTS
Official inference repo for FLUX.1 models
Offline Text To Speech synthesis for python
State-of-the-art TTS model under 25MB
Code for running inference and finetuning with SAM 3 model
Tokenizer-Free TTS for Multilingual Speech Generation
Instant voice cloning by MIT and MyShell. Audio foundation model
Faster Whisper transcription with CTranslate2
Speech-AI-Forge is a project developed around TTS generation model
Robust Speech Recognition via Large-Scale Weak Supervision
A Powerful Native Multimodal Model for Image Generation
A simple native web interface that uses ChatTTS to synthesize text
Use Microsoft Edge's online text-to-speech service from Python
Industrial-level controllable zero-shot text-to-speech system
Open source annotation tool for machine learning practitioners
Nexa SDK is a comprehensive toolkit for supporting ONNX and GGML
Text and image to video generation: CogVideoX and CogVideo
Generate audiobooks from e-books
Ready-to-use OCR with 80+ supported languages
Open image model at the forefront of design
MTEB: Massive Text Embedding Benchmark
Qwen-Image is a powerful image generation foundation model