NVIDIA ParakeetNVIDIA
|
VoxtralMistral AI
|
|||||
Related Products
|
||||||
About
NVIDIA Parakeet-RNNT-1.1B is a multilingual automatic speech recognition model built for quality transcription across voice applications. With 1.1 billion parameters and training on more than 90,000 hours of speech, it supports 25 languages and regional variants, including English, Spanish, French, German, Italian, Arabic, Japanese, Korean, Portuguese, Russian, Hindi, Dutch, Danish, Norwegian, Czech, Polish, Swedish, Thai, Turkish, and Hebrew. The model automatically detects the spoken language and uses a universal tokenizer created by training language-specific tokenizers and merging them into a shared vocabulary, enabling efficient cross-lingual learning and deployment. Parakeet-RNNT produces case-sensitive transcripts with upper and lowercase text, punctuation, spaces, and apostrophes, making the output suitable for production voice applications and downstream language understanding.
|
About
Voxtral models are frontier open source speech‑understanding systems available in two sizes—a 24 B variant for production‑scale applications and a 3 B variant for local and edge deployments, both released under the Apache 2.0 license. They combine high‑accuracy transcription with native semantic understanding, supporting long‑form context (up to 32 K tokens), built‑in Q&A and structured summarization, automatic language detection across major languages, and direct function‑calling to trigger backend workflows from voice. Retaining the text capabilities of their Mistral Small 3.1 backbone, Voxtral handles audio up to 30 minutes for transcription or 40 minutes for understanding and outperforms leading open source and proprietary models on benchmarks such as LibriSpeech, Mozilla Common Voice, and FLEURS. Accessible via download on Hugging Face, API endpoint, or private on‑premises deployment, Voxtral also offers domain‑specific fine‑tuning and advanced enterprise features.
|
|||||
Platforms Supported
Windows
Mac
Linux
Cloud
On-Premises
iPhone
iPad
Android
Chromebook
|
Platforms Supported
Windows
Mac
Linux
Cloud
On-Premises
iPhone
iPad
Android
Chromebook
|
|||||
Audience
Developers and enterprises requiring high-accuracy multilingual automatic speech recognition for transcribing spoken audio across global applications
|
Audience
Developers and product teams requiring a solution to build multilingual voice interfaces and real‑time speech‑to‑action applications
|
|||||
Support
Phone Support
24/7 Live Support
Online
|
Support
Phone Support
24/7 Live Support
Online
|
|||||
API
Offers API
|
API
Offers API
|
|||||
Screenshots and Videos |
Screenshots and Videos |
|||||
Pricing
No information available.
Free Version
Free Trial
|
Pricing
No information available.
Free Version
Free Trial
|
|||||
Reviews/
|
Reviews/
|
|||||
Training
Documentation
Webinars
Live Online
In Person
|
Training
Documentation
Webinars
Live Online
In Person
|
|||||
Company InformationNVIDIA
Founded: 1993
United States
build.nvidia.com/nvidia/parakeet-1_1b-rnnt-multilingual-asr/modelcard
|
Company InformationMistral AI
Founded: 2023
France
mistral.ai/news/voxtral
|
|||||
Alternatives |
Alternatives |
|||||
|
|
|
|||||
|
|
||||||
|
|
|
|||||
|
|
||||||
Categories |
Categories |
|||||
Integrations
ExecuTorch
FluidVoice
Hugging Face
LM Studio Bionic
LazyTyper
Mistral AI
Vision Agents
|
Integrations
ExecuTorch
FluidVoice
Hugging Face
LM Studio Bionic
LazyTyper
Mistral AI
Vision Agents
|
|||||
|
|
|