TextWorld is a sandbox learning environment for the training
ComfyUI wrapper nodes for HunyuanVideo
EPUB to audiobook converter, optimized for Audiobookshelf
Converts text to speech in realtime
State-of-the-art (SoTA) text-to-video pre-trained model
Faster Whisper transcription with CTranslate2
Open-source multi-speaker long-form text-to-speech model
Tokenizer-Free TTS for Multilingual Speech Generation
A robust, efficient, low-latency speech-to-text library
High accuracy RAG for answering questions from scientific documents
Handwritten Text Recognition (HTR) system implemented with TensorFlow
A Family of Open Sourced Music Foundation Models
State-of-the-art TTS model under 25MB
NeuTTS model built from small LLM backbones
AI-powered tool for generating, optimizing, and translating subtitles
Dockerized FastAPI wrapper for Kokoro-82M text-to-speech model
Generate audiobooks from e-books
Official MiniMax Model Context Protocol (MCP) server
Towards Human-Sounding Speech
Powerful Android AI agent with tools, automation, and Linux shell
Windows GUI Automation with Python (based on text properties)
An Open Source text-to-speech system built by inverting Whisper
Video-based AI memory library. Store millions of text chunks in MP4
Qwen3-omni is a natively end-to-end, omni-modal LLM
Google Gen AI Python SDK provides an interface for developers