CosyVoiceAlibaba
|
MAI-Voice-2Microsoft AI
|
|||||
Related Products
|
||||||
About
CosyVoice is Qwen Cloud’s voice cloning and speech synthesis model in the CosyVoice series, designed for professional text-to-speech scenarios with improved sound quality, naturalness, expressiveness, and cloning fidelity. With a short reference recording, it can create a highly similar custom voice without model training; Qwen recommends 10–20 seconds of clear speech, while at least five seconds of continuous speech is required. The model supports real-time, streaming text-to-speech synthesis, allowing applications to accept text and return audio with low first-packet latency. It supports Chinese, English, French, German, Japanese, Korean, and Russian for cloned voices, with language hints available to improve identification during enrollment. Source recordings can use WAV, MP3, or M4A formats and should contain clean speech without background music, noise, or additional speakers.
|
About
MAI-Voice-2 is Microsoft AI’s most expressive and natural-sounding text-to-speech model to date, built for production voice experiences where fidelity, language coverage, speaker consistency, and emotional range directly shape the user experience. It is designed for assistants, customer support, audiobooks, accessibility experiences, games, podcasts, courses, simulations, and creator workflows where voice quality must sound natural, fluid, and trustworthy. It expands from English-only support to 15 languages while maintaining naturalness and expressiveness, with support for English, Italian, French, German, Hindi, Spanish, Portuguese, Korean, Chinese, Turkish, Russian, Thai, Dutch, Romanian, and Hungarian. MAI-Voice-2 offers granular emotion control through tags such as sad, whispered, and excited, along with role-based expressive speech for experiences like motivational trainers, sports commentators, or character voices.
|
|||||
Platforms Supported
Windows
Mac
Linux
Cloud
On-Premises
iPhone
iPad
Android
Chromebook
|
Platforms Supported
Windows
Mac
Linux
Cloud
On-Premises
iPhone
iPad
Android
Chromebook
|
|||||
Audience
Audiobook localization teams that need to clone voices and generate natural multilingual narration at scale
|
Audience
Developers and enterprise product teams that need expressive, multilingual, brand-safe text-to-speech for assistants, customer support, accessibility, education, and long-form audio experiences
|
|||||
Support
Phone Support
24/7 Live Support
Online
|
Support
Phone Support
24/7 Live Support
Online
|
|||||
API
Offers API
|
API
Offers API
|
|||||
Screenshots and Videos |
Screenshots and Videos |
|||||
Pricing
$0.26 per 10,000 characters
Free Version
Free Trial
|
Pricing
No information available.
Free Version
Free Trial
|
|||||
Reviews/
|
Reviews/
|
|||||
Training
Documentation
Webinars
Live Online
In Person
|
Training
Documentation
Webinars
Live Online
In Person
|
|||||
Company InformationAlibaba
Founded: 1999
China
www.qwencloud.com/models/cosyvoice-v3-plus
|
Company InformationMicrosoft AI
Founded: 2024
United States
microsoft.ai/news/mai-voice-2expressive-speech-in-10-languages/
|
|||||
Alternatives |
Alternatives |
|||||
|
|
|
|||||
|
|
|
|||||
|
|
||||||
|
|
||||||
Categories |
Categories |
|||||
|
|
|