Voxtral

Voxtral

Mistral AI
+
+

Related Products

  • Google Cloud Speech-to-Text
    366 Ratings
    Visit Website
  • LTX
    182 Ratings
    Visit Website
  • Iru
    1,382 Ratings
    Visit Website
  • Admin By Request Endpoint Privilege Management
    99 Ratings
    Visit Website
  • QEval
    30 Ratings
    Visit Website
  • Google AI Studio
    30 Ratings
    Visit Website
  • Kitecyber
    64 Ratings
    Visit Website
  • Evertune
    1 Rating
    Visit Website
  • Securden Endpoint Privilege Manager
    7 Ratings
    Visit Website
  • NinjaOne
    6,017 Ratings
    Visit Website

About

Muse Voice Transcribe is Meta’s first real-time audio perception model, delivering streaming automatic speech recognition (ASR), diarization, and endpointing in real time. An autoregressive multimodal model from the Muse Spark family, it processes audio in 80 ms chunks and decides dynamically whether to continue listening or emit text. Its adaptive delay changes the amount of audio context used for each word based on difficulty, balancing transcription accuracy with latency. The model is trained on more than 70 languages, with 25 extensively verified at launch, and natively supports arbitrary code-switching both within and between sentences. Language, keyword, and context biasing can further improve recognition accuracy for specific names, places, contacts, or terminology. Streaming diarization identifies speaker changes and distinguishes more than 20 speakers, while endpointing detects when speech begins and when a user finishes speaking.

About

Voxtral models are frontier open source speech‑understanding systems available in two sizes—a 24 B variant for production‑scale applications and a 3 B variant for local and edge deployments, both released under the Apache 2.0 license. They combine high‑accuracy transcription with native semantic understanding, supporting long‑form context (up to 32 K tokens), built‑in Q&A and structured summarization, automatic language detection across major languages, and direct function‑calling to trigger backend workflows from voice. Retaining the text capabilities of their Mistral Small 3.1 backbone, Voxtral handles audio up to 30 minutes for transcription or 40 minutes for understanding and outperforms leading open source and proprietary models on benchmarks such as LibriSpeech, Mozilla Common Voice, and FLEURS. Accessible via download on Hugging Face, API endpoint, or private on‑premises deployment, Voxtral also offers domain‑specific fine‑tuning and advanced enterprise features.

Platforms Supported

Windows
Mac
Linux
Cloud
On-Premises
iPhone
iPad
Android
Chromebook

Platforms Supported

Windows
Mac
Linux
Cloud
On-Premises
iPhone
iPad
Android
Chromebook

Audience

Developers and AI researchers seeking to build real-time voice applications that transcribe multilingual speech, distinguish speakers, and detect conversational turns

Audience

Developers and product teams requiring a solution to build multilingual voice interfaces and real‑time speech‑to‑action applications

Support

Phone Support
24/7 Live Support
Online

Support

Phone Support
24/7 Live Support
Online

API

Offers API

API

Offers API

Screenshots and Videos

Screenshots and Videos

Pricing

No information available.
Free Version
Free Trial

Pricing

No information available.
Free Version
Free Trial

Reviews/Ratings

Overall 0.0 / 5
ease 0.0 / 5
features 0.0 / 5
design 0.0 / 5
support 0.0 / 5

This software hasn't been reviewed yet. Be the first to provide a review:

Review this Software

Reviews/Ratings

Overall 0.0 / 5
ease 0.0 / 5
features 0.0 / 5
design 0.0 / 5
support 0.0 / 5

This software hasn't been reviewed yet. Be the first to provide a review:

Review this Software

Training

Documentation
Webinars
Live Online
In Person

Training

Documentation
Webinars
Live Online
In Person

Company Information

Meta
Founded: 2004
United States
research.meta.ai/blog/introducing-muse-voice-transcribe

Company Information

Mistral AI
Founded: 2023
France
mistral.ai/news/voxtral

Alternatives

MAI-Transcribe-1.5

MAI-Transcribe-1.5

Microsoft AI

Alternatives

MAI-Transcribe-2

MAI-Transcribe-2

Microsoft AI
Azure AI Speech

Azure AI Speech

Microsoft
Voxtral TTS

Voxtral TTS

Mistral AI

Categories

Categories

Integrations

ExecuTorch
Hugging Face
LM Studio Bionic
LazyTyper
Mistral AI
Vision Agents

Integrations

ExecuTorch
Hugging Face
LM Studio Bionic
LazyTyper
Mistral AI
Vision Agents
Claim Muse Voice Transcribe and update features and information
Claim Muse Voice Transcribe and update features and information
Claim Voxtral and update features and information
Claim Voxtral and update features and information