MAI-Voice-2.1Microsoft
|
Realtime TTS-2Inworld
|
|||||
Related Products
|
||||||
About
MAI-Voice-2.1 is Microsoft's hosted text-to-speech model for developers building voice-enabled products. It generates expressive speech from text, supports multilingual output across 23 languages, and offers emotion and style controls, long-form speaker consistency, and gated matching to consented reference voices. Developers access it through Microsoft Foundry and Azure Speech APIs and SDKs for narration, audiobooks, voice assistants, and customer-service experiences.
|
About
Realtime TTS-2 from Inworld AI is a new generation of voice model built for real-time conversation: a voice model that feels as human as it sounds. It hears the full audio of an exchange, picks up the user’s tone, pacing, and emotional state, then takes voice direction in plain English, the way developers prompt an LLM. Instead of generating speech in isolation, it listens to prior turns of the exchange, so tone and pacing carry forward, and the same line can land differently after a joke than after bad news. Voice Direction lets developers steer delivery like a director would steer a voice actor, using natural-language descriptions rather than fixed emotion presets or sliders. Inline nonverbals like [sigh], [breathe], and [laugh] can be placed inside the text, and the model renders them as audio events. Realtime TTS-2 preserves one voice identity across more than 100 languages, including mid-utterance language switches.
|
|||||
Platforms Supported
Windows
Not Supported
Mac
Not Supported
Linux
Not Supported
Cloud
Supported
On-Premises
Not Supported
iPhone
Not Supported
iPad
Not Supported
Android
Not Supported
Chromebook
Not Supported
|
Platforms Supported
Windows
Not Supported
Mac
Not Supported
Linux
Not Supported
Cloud
Supported
On-Premises
Not Supported
iPhone
Not Supported
iPad
Not Supported
Android
Not Supported
Chromebook
Not Supported
|
|||||
Audience
Developers and organizations building multilingual speech-enabled products, including narration and audiobook services, voice assistants, and customer-service or call-center experiences.
|
Audience
Voice AI developers building realtime agents, characters, tutors, support systems, and companions that need emotionally aware, multilingual, humanlike speech
|
|||||
Support
Phone Support
Not Supported
24/7 Live Support
Not Supported
Online
Not Supported
|
Support
Phone Support
Not Supported
24/7 Live Support
Not Supported
Online
Supported
|
|||||
API
Offers API
Supported
|
API
Offers API
Supported
|
|||||
Screenshots and VideosNo images available
|
Screenshots and Videos |
|||||
Pricing
$22/1M characters
Usage-based at $22 per 1 million characters of generated speech.
Free Version
Not Supported
Free Trial
Not Supported
|
Pricing
$25 per month
Free Version
Supported
Free Trial
Not Supported
|
|||||
Reviews/
|
Reviews/
|
|||||
Training
Documentation
Not Supported
Webinars
Not Supported
Live Online
Not Supported
In Person
Not Supported
|
Training
Documentation
Supported
Webinars
Not Supported
Live Online
Supported
In Person
Not Supported
|
|||||
Company InformationMicrosoft
Founded: 1975
United States
microsoft.ai/models/mai-voice-2-1/
|
Company InformationInworld
Founded: 2021
United States
inworld.ai/blog/realtime-tts-2
|
|||||
Alternatives |
Alternatives |
|||||
|
|
|
|||||
|
|
|
|||||
|
|
|
|||||
|
|
|
|||||
Categories |
Categories |
|||||
Integrations
Azure AI Speech
Supported
ChatGPT
Not Supported
Claude
Not Supported
Gemini
Not Supported
Grok
Not Supported
Perplexity
Not Supported
|
Integrations
Azure AI Speech
Not Supported
ChatGPT
Supported
Claude
Supported
Gemini
Supported
Grok
Supported
Perplexity
Supported
|
|||||
|
|
|