Voxtral

Voxtral

Mistral AI
+
+

Related Products

  • Google Cloud Speech-to-Text
    366 Ratings
    Visit Website
  • LM-Kit.NET
    29 Ratings
    Visit Website
  • Google AI Studio
    41 Ratings
    Visit Website
  • Squaretalk
    300 Ratings
    Visit Website
  • Fathom
    7,733 Ratings
    Visit Website
  • Canopy
    1,025 Ratings
    Visit Website
  • Bigly Sales
    7 Ratings
    Visit Website
  • LTX
    182 Ratings
    Visit Website
  • AdvancedMD
    2 Ratings
    Visit Website
  • Intermedia Unite
    1,631 Ratings
    Visit Website

About

MAI-Transcribe-2-Streaming is a low-latency streaming transcription model built for real-time speech applications, delivering transcripts in 60 languages with automatic, continuous language detection. Rather than waiting for someone to finish speaking, it produces initial partial transcripts just over 100 ms after receiving audio, revises them as more context arrives, and commits stable text quickly. This allows voice applications to begin reasoning, calling tools, or displaying live transcripts while a person is still speaking. Microsoft reports that the model ranks No. 1 for both final and partial transcript accuracy on Artificial Analysis. MAI-Voice-2.1 complements it with multilingual text-to-speech across 23 languages and 26 locales, allowing a single voice to switch languages while maintaining the same speaker identity and adopting native accents.

About

Voxtral models are frontier open source speech‑understanding systems available in two sizes—a 24 B variant for production‑scale applications and a 3 B variant for local and edge deployments, both released under the Apache 2.0 license. They combine high‑accuracy transcription with native semantic understanding, supporting long‑form context (up to 32 K tokens), built‑in Q&A and structured summarization, automatic language detection across major languages, and direct function‑calling to trigger backend workflows from voice. Retaining the text capabilities of their Mistral Small 3.1 backbone, Voxtral handles audio up to 30 minutes for transcription or 40 minutes for understanding and outperforms leading open source and proprietary models on benchmarks such as LibriSpeech, Mozilla Common Voice, and FLEURS. Accessible via download on Hugging Face, API endpoint, or private on‑premises deployment, Voxtral also offers domain‑specific fine‑tuning and advanced enterprise features.

Platforms Supported

Windows Not Supported
Mac Not Supported
Linux Not Supported
Cloud Supported
On-Premises Not Supported
iPhone Not Supported
iPad Not Supported
Android Not Supported
Chromebook Not Supported

Platforms Supported

Windows Not Supported
Mac Not Supported
Linux Not Supported
Cloud Supported
On-Premises Not Supported
iPhone Not Supported
iPad Not Supported
Android Not Supported
Chromebook Not Supported

Audience

Developers and AI teams seeking to build fast, multilingual voice agents, real-time transcription systems, and conversational speech applications

Audience

Developers and product teams requiring a solution to build multilingual voice interfaces and real‑time speech‑to‑action applications

Support

Phone Support Not Supported
24/7 Live Support Not Supported
Online Supported

Support

Phone Support Supported
24/7 Live Support Not Supported
Online Supported

API

Offers API Not Supported

API

Offers API Supported

Screenshots and Videos

Screenshots and Videos

Pricing

No information available.
Free Version Not Supported
Free Trial Not Supported

Pricing

No information available.
Free Version Not Supported
Free Trial Not Supported

Reviews/Ratings

Overall 0.0 / 5
ease 0.0 / 5
features 0.0 / 5
design 0.0 / 5
support 0.0 / 5

This software hasn't been reviewed yet. Be the first to provide a review:

Review this Software

Reviews/Ratings

Overall 0.0 / 5
ease 0.0 / 5
features 0.0 / 5
design 0.0 / 5
support 0.0 / 5

This software hasn't been reviewed yet. Be the first to provide a review:

Review this Software

Training

Documentation Supported
Webinars Not Supported
Live Online Not Supported
In Person Not Supported

Training

Documentation Supported
Webinars Not Supported
Live Online Supported
In Person Supported

Company Information

Microsoft AI
Founded: 2024
United States
microsoft.ai/news/our-first-streaming-transcription-model/

Company Information

Mistral AI
Founded: 2023
France
mistral.ai/news/voxtral

Alternatives

Alternatives

Cartesia Ink 2

Cartesia Ink 2

Cartesia
Azure AI Speech

Azure AI Speech

Microsoft

Categories

AI Models Supported
Speech to Text Supported

Categories

AI Models Supported
Speech to Text Supported
Transcription Supported

Integrations

ExecuTorch Not Supported
Hugging Face Not Supported
LM Studio Bionic Not Supported
LazyTyper Not Supported
Mistral AI Not Supported
Vision Agents Not Supported

Integrations

ExecuTorch Supported
Hugging Face Supported
LM Studio Bionic Supported
LazyTyper Supported
Mistral AI Supported
Vision Agents Supported
Claim MAI-Transcribe-2-Streaming and update features and information
Claim MAI-Transcribe-2-Streaming and update features and information
Claim Voxtral and update features and information
Claim Voxtral and update features and information