MAI-Voice-2.1Microsoft
|
Orpheus TTSCanopy Labs
|
|||||
Related Products
|
||||||
About
MAI-Voice-2.1 is Microsoft's hosted text-to-speech model for developers building voice-enabled products. It generates expressive speech from text, supports multilingual output across 23 languages, and offers emotion and style controls, long-form speaker consistency, and gated matching to consented reference voices. Developers access it through Microsoft Foundry and Azure Speech APIs and SDKs for narration, audiobooks, voice assistants, and customer-service experiences.
|
About
Canopy Labs has introduced Orpheus, a family of state-of-the-art speech large language models (LLMs) designed for human-level speech generation. These models are built on the Llama-3 architecture and are trained on over 100,000 hours of English speech data, enabling them to produce natural intonation, emotion, and rhythm that surpasses current state-of-the-art closed source models. Orpheus supports zero-shot voice cloning, allowing users to replicate voices without prior fine-tuning, and offers guided emotion and intonation control through simple tags. The models achieve low latency, with approximately 200ms streaming latency for real-time applications, reducible to around 100ms with input streaming. Canopy Labs has released both pre-trained and fine-tuned 3B-parameter models under the permissive Apache 2.0 license, with plans to release smaller models of 1B, 400M, and 150M parameters for use on resource-constrained devices.
|
|||||
Platforms Supported
Windows
Not Supported
Mac
Not Supported
Linux
Not Supported
Cloud
Supported
On-Premises
Not Supported
iPhone
Not Supported
iPad
Not Supported
Android
Not Supported
Chromebook
Not Supported
|
Platforms Supported
Windows
Not Supported
Mac
Not Supported
Linux
Not Supported
Cloud
Supported
On-Premises
Not Supported
iPhone
Not Supported
iPad
Not Supported
Android
Not Supported
Chromebook
Not Supported
|
|||||
Audience
Developers and organizations building multilingual speech-enabled products, including narration and audiobook services, voice assistants, and customer-service or call-center experiences.
|
Audience
Researchers needing a solution offering high-quality, low-latency speech synthesis with customizable voice cloning and emotion control capabilities
|
|||||
Support
Phone Support
Not Supported
24/7 Live Support
Not Supported
Online
Not Supported
|
Support
Phone Support
Not Supported
24/7 Live Support
Not Supported
Online
Supported
|
|||||
API
Offers API
Supported
|
API
Offers API
Not Supported
|
|||||
Screenshots and VideosNo images available
|
Screenshots and Videos |
|||||
Pricing
$22/1M characters
Usage-based at $22 per 1 million characters of generated speech.
Free Version
Not Supported
Free Trial
Not Supported
|
Pricing
No information available.
Free Version
Not Supported
Free Trial
Supported
|
|||||
Reviews/
|
Reviews/
|
|||||
Training
Documentation
Not Supported
Webinars
Not Supported
Live Online
Not Supported
In Person
Not Supported
|
Training
Documentation
Supported
Webinars
Not Supported
Live Online
Supported
In Person
Not Supported
|
|||||
Company InformationMicrosoft
Founded: 1975
United States
microsoft.ai/models/mai-voice-2-1/
|
Company InformationCanopy Labs
United States
canopylabs.ai/model-releases
|
|||||
Alternatives |
Alternatives |
|||||
|
|
|
|||||
|
|
|
|||||
|
|
|
|||||
|
|
|
|||||
Categories |
Categories |
|||||
Integrations
Azure AI Speech
Supported
Baseten
Not Supported
GitHub
Not Supported
Google Colab
Not Supported
Hugging Face
Not Supported
Llama 3
Not Supported
VoiSpark
Not Supported
|
Integrations
Azure AI Speech
Not Supported
Baseten
Supported
GitHub
Supported
Google Colab
Supported
Hugging Face
Supported
Llama 3
Supported
VoiSpark
Supported
|
|||||
|
|
|