Fast multimodal LLM for real-time voice interaction and AI apps
Repo of Qwen2-Audio chat & pretrained large audio language model
A specialized Claude Code workspace for creating long-form
Curated collection of Amazing Python scripts
Toolkit for conversational AI
Autonomous agents for everyone
AI suite powered by state-of-the-art models and providing advanced AI
Where Models and Agents Co-Evolve
A GPT-4o Level MLLM for Vision, Speech and Multimodal Live Streaming
A natural language interface for computers
A Claude skill that automatically posts personalized comments
HTML5 js recording mp3 wav ogg webm amr format
Convert VoIP calls to text and analyze them with AI
Live analysis of pitches, harmonics, chords, and keys.
Application which detects musical notes from the microphone.
Virtual AI anchor that combines state-of-the-art technology
Chat & pretrained large audio language model proposed by Alibaba Cloud
Toolkit for audio, music, and speech generation
2D open source actuator simulation software
3D open source actuator simulation software
Voice dialogue, role-playing, multi-topic discussion, picture creation
General Speech Restoration
A python package to analyze and compare voices with deep learning
Code for the Psygraph mobile application