CosyVoiceAlibaba
|
Gemini Live APIGoogle
|
|||||
Related Products
|
||||||
About
CosyVoice is Qwen Cloud’s voice cloning and speech synthesis model in the CosyVoice series, designed for professional text-to-speech scenarios with improved sound quality, naturalness, expressiveness, and cloning fidelity. With a short reference recording, it can create a highly similar custom voice without model training; Qwen recommends 10–20 seconds of clear speech, while at least five seconds of continuous speech is required. The model supports real-time, streaming text-to-speech synthesis, allowing applications to accept text and return audio with low first-packet latency. It supports Chinese, English, French, German, Japanese, Korean, and Russian for cloned voices, with language hints available to improve identification during enrollment. Source recordings can use WAV, MP3, or M4A formats and should contain clean speech without background music, noise, or additional speakers.
|
About
The Gemini Live API is a preview feature that enables low-latency, bidirectional voice and video interactions with Gemini. It allows end users to experience natural, human-like voice conversations and provides the ability to interrupt the model's responses using voice commands. The model can process text, audio, and video input, and it can provide text and audio output. New capabilities include two new voices and 30 new languages with configurable output language, configurable image resolutions (66/256 tokens), configurable turn coverage (send all inputs all the time or only when the user is speaking), configurable interruption settings, configurable voice activity detection, new client events for end-of-turn signaling, token counts, a client event for signaling the end of stream, text streaming, configurable session resumption with session data stored on the server for 24 hours, and longer session support with a sliding context window.
|
|||||
Platforms Supported
Windows
Mac
Linux
Cloud
On-Premises
iPhone
iPad
Android
Chromebook
|
Platforms Supported
Windows
Mac
Linux
Cloud
On-Premises
iPhone
iPad
Android
Chromebook
|
|||||
Audience
Users, developers and teams that need to clone voices and generate natural multilingual narration at scale
|
Audience
Researchers looking for a solution to build real-time, multimodal AI applications that require low-latency voice and video interactions
|
|||||
Support
Phone Support
24/7 Live Support
Online
|
Support
Phone Support
24/7 Live Support
Online
|
|||||
API
Offers API
|
API
Offers API
|
|||||
Screenshots and Videos |
Screenshots and Videos |
|||||
Pricing
$0.26 per 10,000 characters
Free Version
Free Trial
|
Pricing
No information available.
Free Version
Free Trial
|
|||||
Reviews/
|
Reviews/
|
|||||
Training
Documentation
Webinars
Live Online
In Person
|
Training
Documentation
Webinars
Live Online
In Person
|
|||||
Company InformationAlibaba
Founded: 1999
China
www.qwencloud.com/models/cosyvoice-v3-plus
|
Company InformationGoogle
Founded: 1998
United States
ai.google.dev/gemini-api/docs/live
|
|||||
Alternatives |
Alternatives |
|||||
|
|
||||||
|
|
|
|||||
|
|
|
|||||
|
|
||||||
Categories |
Categories |
|||||
Integrations
Daily
Firebase
Fishjam
Gemini 3 Pro Image
Gemini 3.1 Flash Image
Gemini 3.1 Flash Live
Gemini 3.5 Live Translate
Gemini Enterprise
Gemini Enterprise Agent Platform
Google AI Studio
|
Integrations
Daily
Firebase
Fishjam
Gemini 3 Pro Image
Gemini 3.1 Flash Image
Gemini 3.1 Flash Live
Gemini 3.5 Live Translate
Gemini Enterprise
Gemini Enterprise Agent Platform
Google AI Studio
|
|||||
|
|
|