Video-based AI memory library. Store millions of text chunks in MP4
State-of-the-art Machine Learning for Pytorch, TensorFlow, and JAX
SoTA open-source TTS
Google Gen AI Python SDK provides an interface for developers
Ready-to-use OCR with 80+ supported languages
Industrial-level controllable zero-shot text-to-speech system
Faster Whisper transcription with CTranslate2
Library for OCR-related tasks powered by Deep Learning
Wan2.2: Open and Advanced Large-Scale Video Generative Model
The simplest, fastest repository for training/finetuning models
Wan2.1: Open and Advanced Large-Scale Video Generative Model
A Family of Open Sourced Music Foundation Models
Open source healthcare AI
Stanford NLP Python library for many human languages
Open-source multi-speaker long-form text-to-speech model
Official inference repo for FLUX.2 models
ComfyUI integration for Microsoft's VibeVoice text-to-speech model
Clone a voice in 5 seconds to generate arbitrary speech in real-time
Open-source image generative foundation model
A high-quality rapid TTS voice cloning model
Handwritten Text Recognition (HTR) system implemented with TensorFlow
Official MiniMax Model Context Protocol (MCP) server
Converts text to speech in realtime
Generate audiobooks from e-books
AI-powered tool for generating, optimizing, and translating subtitles