VoxtralMistral AI
|
||||||
Related Products
|
||||||
About
Deploy accurate speech recognition at scale while continuously improving model performance by labeling data and training from a single console. We deliver state-of-the-art speech recognition and understanding at scale. We do it by providing cutting-edge model training and data-labeling alongside flexible deployment options. Our platform recognizes multiple languages, accents, and words, dynamically tuning to the needs of your business with every training session. The fastest, most accurate, most reliable, most scalable speech transcription, with understanding — rebuilt just for enterprise. We’ve reinvented ASR with 100% deep learning that allows companies to continuously improve accuracy. Stop waiting for the big tech players to improve their software and forcing your developers to manually boost accuracy with keywords in every API call. Start training your speech model and reaping the benefits in weeks, not months or years.
|
About
Voxtral models are frontier open source speech‑understanding systems available in two sizes—a 24 B variant for production‑scale applications and a 3 B variant for local and edge deployments, both released under the Apache 2.0 license. They combine high‑accuracy transcription with native semantic understanding, supporting long‑form context (up to 32 K tokens), built‑in Q&A and structured summarization, automatic language detection across major languages, and direct function‑calling to trigger backend workflows from voice. Retaining the text capabilities of their Mistral Small 3.1 backbone, Voxtral handles audio up to 30 minutes for transcription or 40 minutes for understanding and outperforms leading open source and proprietary models on benchmarks such as LibriSpeech, Mozilla Common Voice, and FLEURS. Accessible via download on Hugging Face, API endpoint, or private on‑premises deployment, Voxtral also offers domain‑specific fine‑tuning and advanced enterprise features.
|
|||||
Platforms Supported
Windows
Not Supported
Mac
Not Supported
Linux
Not Supported
Cloud
Supported
On-Premises
Not Supported
iPhone
Not Supported
iPad
Not Supported
Android
Not Supported
Chromebook
Not Supported
|
Platforms Supported
Windows
Not Supported
Mac
Not Supported
Linux
Not Supported
Cloud
Supported
On-Premises
Not Supported
iPhone
Not Supported
iPad
Not Supported
Android
Not Supported
Chromebook
Not Supported
|
|||||
Audience
Companies looking for Speech to Text (STT) API for real-time and batch transcriptions, on premise or in the cloud.
|
Audience
Developers and product teams requiring a solution to build multilingual voice interfaces and real‑time speech‑to‑action applications
|
|||||
Support
Phone Support
Supported
24/7 Live Support
Not Supported
Online
Not Supported
|
Support
Phone Support
Supported
24/7 Live Support
Not Supported
Online
Supported
|
|||||
API
Offers API
Supported
|
API
Offers API
Supported
|
|||||
Screenshots and Videos |
Screenshots and Videos |
|||||
Pricing
$0
Free Version
Supported
Free Trial
Supported
|
Pricing
No information available.
Free Version
Not Supported
Free Trial
Not Supported
|
|||||
Reviews/
|
Reviews/
|
|||||
Training
Documentation
Supported
Webinars
Not Supported
Live Online
Not Supported
In Person
Not Supported
|
Training
Documentation
Supported
Webinars
Not Supported
Live Online
Supported
In Person
Supported
|
|||||
Company InformationDeepgram
Founded: 2015
United States
deepgram.com
|
Company InformationMistral AI
Founded: 2023
France
mistral.ai/news/voxtral
|
|||||
Alternatives |
Alternatives |
|||||
|
|
||||||
|
|
||||||
|
|
|
|||||
|
|
||||||
Categories |
Categories |
|||||
Transcription Features
AI / Machine Learning
Supported
Annotations
Not Supported
Audio/Video File Upload
Supported
Automatic Transcription
Supported
Collaboration Tools
Not Supported
File Sharing
Not Supported
For Manual Transcription
Not Supported
Full Text Search
Not Supported
Multi-Language Support
Supported
Natural Language Processing (NLP)
Supported
Playback Controls
Not Supported
Speech Recognition
Supported
Subtitles
Not Supported
Text Editor
Not Supported
Timecoding
Not Supported
Speech Recognition Features
Audio Capture
Not Supported
Automatic Form Fill
Not Supported
Automatic Transcription
Supported
Call Analysis
Not Supported
Concatenated Speech
Not Supported
Continuous Speech
Not Supported
Customizable Macros
Not Supported
Multi-Languages
Supported
Specialty Vocabularies
Supported
Speech-to-Text Analysis
Not Supported
Variable Frequency
Not Supported
Voice Recognition
Supported
|
||||||
Integrations
Vision Agents
Supported
Agentic.Market
Supported
Creovai
Supported
Deepgram Saga
Supported
Fluents.ai
Supported
Genesys Cloud CX
Supported
Hugging Face
Not Supported
Hunch
Supported
Kubernetes
Supported
LazyTyper
Not Supported
|
Integrations
Vision Agents
Supported
Agentic.Market
Not Supported
Creovai
Not Supported
Deepgram Saga
Not Supported
Fluents.ai
Not Supported
Genesys Cloud CX
Not Supported
Hugging Face
Supported
Hunch
Not Supported
Kubernetes
Not Supported
LazyTyper
Supported
|
|||||
|
|
|