Google Cloud’s Speech API processes more than 1 billion voice minutes per month with close to human levels of understanding for many commonly spoken languages. Powered by the best of Google's AI research and technology, Google Cloud's Speech-to-Text API helps you accurately transcribe speech into text in 73 languages and 137 different local variants. Leverage Google’s most advanced deep learning neural network algorithms for automatic speech recognition (ASR) and deploy ASR wherever you need it, whether in the cloud with the API, on-premises with Speech-to-Text On-Prem, or locally on any device with Speech On-Device.
Learn more
LM-Kit.NET is a complete local AI runtime for .NET that lets engineering teams ship AI-powered features without cloud dependencies, per-token costs, or data leaving the network.
Most .NET AI integrations stop at inference. LM-Kit.NET covers the full range of capabilities production applications actually need: agentic workflows with tool calling, planning, and memory; document intelligence with OCR and structured extraction; retrieval-augmented generation with built-in vector storage; multilingual speech-to-text; vision and multimodal understanding; text analysis with classification, NER, PII extraction, and sentiment; and text generation with translation, summarization, and constrained output.
Ships in one NuGet package, runs in-process with no sidecar services, and works across all major hardware acceleration backends. Drop-in replacement for Semantic Kernel through its Microsoft.Extensions.AI compatibility layer.
Learn more
Spokenly
Spokenly is an AI-powered dictation app for Mac, iPhone, Windows, and Linux that turns speech into clean, punctuated text wherever you work. Hold a shortcut, speak naturally, and release to place the transcription directly at the cursor in browsers, email, chat, word processors, IDEs, terminals, and other apps. It supports more than 100 languages, including mixed-language dictation, and offers both local and cloud speech-to-text models. Whisper, Parakeet, and other on-device models can run completely offline, while cloud engines from providers such as OpenAI, Deepgram, Groq, Soniox, and ElevenLabs can be used for higher-accuracy or real-time transcription. Local Only Mode blocks network requests so voice data stays on the device. Modes let users save different transcription models, AI providers, prompts, and output styles for specific tasks, while AI Instructions can remove filler words, fix grammar and punctuation, summarize, rewrite, translate, or reformat dictated text.
Learn more