Riverside
Riverside (previously "Riverside FM") is an all-in-one AI-powered content creation studio for recording, editing, and streaming high-quality video and audio. Designed for podcasters, marketers, and businesses, Riverside captures 4K video and lossless audio locally for every participant—ensuring crystal-clear quality even with weak connections. Its intuitive text-based editor lets users trim, clean up, and caption recordings directly from the transcript, eliminating the need for complex editing tools. With features like Magic Audio, AI Voice, and VideoDub, creators can polish sound, fix mistakes, and sync lips with AI-generated speech in seconds. Riverside also enables HD live streaming and AI Show Notes for automatic titles, chapters, and keywords that simplify publishing. Whether recording a podcast, webinar, or social clip, Riverside brings professional-grade production within everyone’s reach.
Learn more
Grok Speech to Text (STT)
Grok Speech to Text is a standalone audio API built to help developers integrate fast, accurate transcription into any application. Built on the same stack that powers Grok Voice, Tesla vehicles, and Starlink customer support, the API is designed for use cases such as voice agents, real-time transcription tools, accessibility solutions, podcasts, meeting capture, telephony, and interactive audio experiences. Grok STT can generate transcripts from large audio files through a REST API or transcribe speech in real time through a low-latency WebSocket API. It includes word-level timestamps, speaker diarization, multichannel support, and intelligent Inverse Text Normalization that converts spoken language into properly formatted structured output for numbers, dates, currencies, and more. Grok Speech to Text is evaluated across phone calls, meetings, video and podcast content, and telephony, with strong performance in entity recognition and business use cases.
Learn more
Pepys
Pepys is pay-as-you-go AI transcription software for turning audio and video into speaker-labelled, timestamped transcripts. It supports multilingual transcription, AI-powered transcript search and chat, summaries, translation, exports, a developer API and MCP access.
Upload a file or paste a link—from YouTube, TikTok, Instagram, Facebook, Spotify, or Apple Podcasts—and get a clean transcript with word and segment-level timestamps plus speaker labels. Exports: TXT, Markdown, DOCX, PDF, SRT, VTT, JSON.
Learn more
MAI-Transcribe-2
MAI-Transcribe-2 is Microsoft AI’s most capable transcription model yet, designed to deliver fast, accurate speech recognition across a broad range of real-world audio. It supports speaker diarization to distinguish speakers and attribute words to the right person, along with word-level timestamps for precise alignment, search, navigation, and editing. Keyword biasing helps recognize domain-specific terminology, abbreviations, names, and other terms that can be difficult to distinguish from context alone. Developers can choose between configurable transcription styles: a verbatim setting that preserves filler words and false starts for compliance and analysis, or a clean setting that removes fillers for more readable captions, notes, and published transcripts. The model supports code-switching for conversations that naturally move between languages, including blended language pairs such as Hinglish and Spanglish, and can automatically identify the language being spoken.
Learn more