Official MiniMax Model Context Protocol (MCP) server
Persian NLP Toolkit
Framework for building real-time voice and multimodal AI agents
Controllable and fast Text-to-Speech for over 7000 languages
Have a natural, spoken conversation with AI
SOTA discrete acoustic codec models with 40/75 tokens per second
Audio foundation model excelling in audio understanding
A GPT-4o Level MLLM for Vision, Speech and Multimodal Live Streaming
Real-time voice interactive digital human
Style-Bert-VITS2: Bert-VITS2 with more controllable voice styles
Clone a voice in 5 seconds to generate arbitrary speech in real-time
EPUB to audiobook converter, optimized for Audiobookshelf
Powerful Android AI agent with tools, automation, and Linux shell
Management of Yandex Station and other smart home devices
A 0.1B Omni model trained from scratch
NLP Cloud serves high performance pre-trained or custom models for NER
Open-source industrial-grade ASR models
Framework for building neural networks
Underthesea - Vietnamese NLP Toolkit
Repo of Qwen2-Audio chat & pretrained large audio language model
FAIR Sequence Modeling Toolkit 2
Official PyTorch Implementation
Han Language Processing
Bailing is a voice dialogue robot similar to GPT-4o
Open source AI VTuber platform with voice chat and Live2D avatars