Ready-to-use OCR with 80+ supported languages
SoTA open-source TTS
Converts text to speech in realtime
Google Gen AI Python SDK provides an interface for developers
Tokenizer-Free TTS for Multilingual Speech Generation
Industrial-level controllable zero-shot text-to-speech system
Faster Whisper transcription with CTranslate2
Wan2.2: Open and Advanced Large-Scale Video Generative Model
Library for OCR-related tasks powered by Deep Learning
A Family of Open Sourced Music Foundation Models
Stanford NLP Python library for many human languages
Wan2.1: Open and Advanced Large-Scale Video Generative Model
The simplest, fastest repository for training/finetuning models
Style-Bert-VITS2: Bert-VITS2 with more controllable voice styles
Voice Recognition to Text Tool
Open source healthcare AI
ComfyUI integration for Microsoft's VibeVoice text-to-speech model
Open-source image generative foundation model
Clone a voice in 5 seconds to generate arbitrary speech in real-time
Official inference repo for FLUX.2 models
Generate audiobooks from e-books, voice cloning & 1107+ languages
Easily compute clip embeddings and build a clip retrieval system
Easy-to-use and powerful NLP library with Awesome model zoo
Handwritten Text Recognition (HTR) system implemented with TensorFlow
Official MiniMax Model Context Protocol (MCP) server