Qwen-Audio-3.0-TTS-Plus
Qwen-Audio-3.0-TTS-Plus is the high-quality variant of Qwen-Audio-3.0-TTS, optimized for naturalness and timbre fidelity when output quality matters more than speed. It supports 16 languages, plus improved fidelity for several Chinese dialects. The model delivers strong multilingual intelligibility and ranks first in speaker similarity across all supported languages, helping cloned voices remain recognizable and consistent across linguistic contexts. Developers can direct delivery through ordinary natural-language instructions instead of manually tuning acoustic parameters, controlling emotion, role, scenario, pacing, projection, and tone with simple prompts. Inline tags provide fine-grained control over breaths, laughter, emotional shifts, and other non-verbal details, making the model useful for narration, games, character dialogue, and dubbing.
Learn more
Grok Text to Speech (TTS)
Grok Text to Speech (TTS) is a standalone audio API built to help developers generate fast, natural, and expressive speech from text. Built on the same stack that powers Grok Voice, Tesla vehicles, and Starlink customer support, the API makes it straightforward to integrate high-quality voice generation into applications such as voice agents, accessibility tools, podcasts, assistants, customer experiences, and interactive audio products. Grok TTS can turn long-form text into speech through a REST API or generate speech in real time through a WebSocket API, giving developers flexibility for both batch audio generation and live conversational experiences. It is designed around expressive delivery, not just flat narration, with fine-grained control through simple inline and wrapping speech tags. Developers can add natural prosody and emotion using tags, allowing lifelike delivery without complex markup.
Learn more
Qwen-Audio-3.0-TTS-Flash
Qwen-Audio-3.0-TTS-Flash is the real-time variant of Qwen-Audio-3.0-TTS, tuned for interactive applications with first-packet latency at the 300 ms level. It supports 16 languages, along with improved fidelity for several Chinese dialects. Across multilingual evaluations, Flash delivers the lowest average WER/CER in the family at 3.87, showing strong intelligibility while preserving speaker identity across diverse languages. Developers can guide delivery with plain-language instructions instead of manually adjusting acoustic parameters, controlling emotion, role, scenario, pace, projection, and tone through simple prompts. Inline tags add precise non-verbal details, making the model well-suited to conversational agents, narration, games, dubbing, and other expressive speech experiences. Voice cloning is designed to work with imperfect reference audio; targeted acoustic simulation suppresses noise and reverberation while retaining the original speaker’s timbre.
Learn more
Cartesia Sonic-3.5
Sonic 3.5 is Cartesia’s fastest, most natural text-to-speech model, built for expressive, real-time voice generation with sub-90ms latency and native support for 42 languages. It is designed to follow transcripts faithfully, voice confirmation codes, and heteronyms correctly without preprocessing, and stay expressive enough to carry a real conversation. It supports languages intended to deliver native-quality speech. Sonic 3.5 focuses on clean audio across every language and voice, with no artifacts to edit out, making it practical for production voice experiences where quality, speed, and consistency matter. Its expressive conversational delivery provides strong pacing and real emotional range, tuned for support and agent transcripts. Alphanumerics such as order numbers, phone numbers, IDs, and emails are spoken naturally in every language, while context-aware English pronunciation helps words like read, bass, and bow land correctly from the surrounding text.
Learn more