Instant voice cloning by MIT and MyShell. Audio foundation model
A Powerful Native Multimodal Model for Image Generation
An AI-agent skill that turns Markdown into paste-ready WeChat article
Official inference repo for FLUX.2 models
Open image model at the forefront of design
Library for OCR-related tasks powered by Deep Learning
Ready-to-use OCR with 80+ supported languages
A simple native web interface that uses ChatTTS to synthesize text
Voice Recognition to Text Tool
Self-host the powerful Chatterbox TTS model
A robust, efficient, low-latency speech-to-text library
Generate audiobooks from e-books, voice cloning & 1107+ languages
Qwen-Image is a powerful image generation foundation model
EPUB to audiobook converter, optimized for Audiobookshelf
Text and image to video generation: CogVideoX and CogVideo
Qwen3-omni is a natively end-to-end, omni-modal LLM
Easy-to-use and powerful NLP library with Awesome model zoo
A simple, high-quality voice conversion tool focused on ease of use
MTEB: Massive Text Embedding Benchmark
Run Bonsai (1-bit) and Ternary-Bonsai language models locally
Framework for building realtime multimodal voice AI agents apps
A text-to-speech, speech-to-text and speech-to-speech library
Python & command-line tool to gather text on the Web
State-of-the-art Machine Learning for Pytorch, TensorFlow, and JAX
An Open Source text-to-speech system built by inverting Whisper