Convert Google Gemini web into OpenAI-compatible API
Long-form streaming TTS system for multi-speaker dialogue generation
MOSS‑TTS Family open‑source speech and sound generation model
A GPT-4o Level MLLM for Vision, Speech and Multimodal Live Streaming
Provides convenient access to the Anthropic REST API from any Python 3
Capable of understanding text, audio, vision, video
A 0.1B Omni model trained from scratch
Qwen3-omni is a natively end-to-end, omni-modal LLM
Python SDK for Claude Agent
Controllable & emotion-expressive zero-shot TTS
FAIR Sequence Modeling Toolkit 2
ChatGPT interface with better UI
Code for the paper Hybrid Spectrogram and Waveform Source Separation