A Powerful Native Multimodal Model for Image Generation
ComfyUI integration for Microsoft's VibeVoice text-to-speech model
Management of Yandex Station and other smart home devices
EPUB to audiobook converter, optimized for Audiobookshelf
A robust, efficient, low-latency speech-to-text library
Label Studio is a multi-type data labeling and annotation tool
Handwritten Text Recognition (HTR) system implemented with TensorFlow
An open-source toolkit for monitoring Language Learning Models (LLMs)
Run Bonsai (1-bit) and Ternary-Bonsai language models locally
Framework for building realtime multimodal voice AI agents apps
An Open Source text-to-speech system built by inverting Whisper
Voice Recognition to Text Tool
A simple native web interface that uses ChatTTS to synthesize text
Qwen-Image is a powerful image generation foundation model
Text and image to video generation: CogVideoX and CogVideo
Implementation of Imagen, Google's Text-to-Image Neural Network
Easy-to-use and powerful NLP library with Awesome model zoo
A simple, high-quality voice conversion tool focused on ease of use
Self-host the powerful Chatterbox TTS model
Powerful Android AI agent with tools, automation, and Linux shell
Open-source multi-speaker long-form text-to-speech model
Dockerized FastAPI wrapper for Kokoro-82M text-to-speech model
State-of-the-art Machine Learning for Pytorch, TensorFlow, and JAX
Free, high-quality text-to-speech API endpoint to replace OpenAI
Agent harness to make your slop code well-engineered and beautiful