Inworld Realtime STTInworld
|
||||||
Related Products
|
||||||
About
Inworld Realtime STT is a realtime streaming STT API that understands users beyond their words. It combines low-latency speech recognition with voice profiling, extracting emotion, vocal style, accent, age, and pitch directly from raw audio so downstream LLMs and TTS systems can respond with more adaptive, expressive behavior. Developers can stream audio in real time, transcribe complete files, or extract voice profile signals through one unified API, with realtime bidirectional streaming over WebSocket, synchronous transcription for full audio files, voice profile signals on every streaming chunk, and multi-provider support through a single model ID. Every audio chunk can produce a realtime profile of the speaker with confidence scores, giving LLMs structured context such as whether a user sounds sad, frustrated, soft, high-pitched, or calm.
|
About
Muse Voice Transcribe is Meta’s first real-time audio perception model, delivering streaming automatic speech recognition (ASR), diarization, and endpointing in real time. An autoregressive multimodal model from the Muse Spark family, it processes audio in 80 ms chunks and decides dynamically whether to continue listening or emit text. Its adaptive delay changes the amount of audio context used for each word based on difficulty, balancing transcription accuracy with latency. The model is trained on more than 70 languages, with 25 extensively verified at launch, and natively supports arbitrary code-switching both within and between sentences. Language, keyword, and context biasing can further improve recognition accuracy for specific names, places, contacts, or terminology. Streaming diarization identifies speaker changes and distinguishes more than 20 speakers, while endpointing detects when speech begins and when a user finishes speaking.
|
|||||
Platforms Supported
Windows
Mac
Linux
Cloud
On-Premises
iPhone
iPad
Android
Chromebook
|
Platforms Supported
Windows
Mac
Linux
Cloud
On-Premises
iPhone
iPad
Android
Chromebook
|
|||||
Audience
AI teams building realtime assistants that need fast transcription, speaker context, multilingual support, and emotionally adaptive responses
|
Audience
Developers and AI researchers seeking to build real-time voice applications that transcribe multilingual speech, distinguish speakers, and detect conversational turns
|
|||||
Support
Phone Support
24/7 Live Support
Online
|
Support
Phone Support
24/7 Live Support
Online
|
|||||
API
Offers API
|
API
Offers API
|
|||||
Screenshots and Videos |
Screenshots and Videos |
|||||
Pricing
Free
Free Version
Free Trial
|
Pricing
No information available.
Free Version
Free Trial
|
|||||
Reviews/
|
Reviews/
|
|||||
Training
Documentation
Webinars
Live Online
In Person
|
Training
Documentation
Webinars
Live Online
In Person
|
|||||
Company InformationInworld
Founded: 2021
United States
inworld.ai/speech-to-text
|
Company InformationMeta
Founded: 2004
United States
research.meta.ai/blog/introducing-muse-voice-transcribe
|
|||||
Alternatives |
Alternatives |
|||||
|
|
|
|||||
|
|
|
|||||
|
|
|
|||||
|
|
|
|||||
Categories |
Categories |
|||||
Integrations
No info available.
|
Integrations
No info available.
|
|||||
|
|
|