Silkwave Voice
Silkwave Voice is a privacy-focused audio recording and transcription app for macOS. Record from your microphone, system audio, or both at once - with accurate, real-time transcription powered by Apple's on-device speech-to-text models. No cloud uploads, no subscriptions, no per-minute API costs.
RECORD ANY AUDIO SOURCE
• Microphone - voice notes, in-person meetings, dictation
• System Audio - Zoom, Google Meet, Teams, YouTube, browser tabs
• Both at once - capture your mic and remote participants simultaneously
ON-DEVICE TRANSCRIPTION
• Real-time speech-to-text using Apple's on-device models
• 10 languages: Cantonese, Chinese, English, French, German, Italian, Japanese, Korean, Portuguese, Spanish
• Completely local - no internet connection needed
AI-POWERED SUMMARIES
• Structured summaries with key topics, action items, and decisions
• Powered by ChatGPT through Apple Intelligence - no API keys needed
Learn more
Zeemo AI
Simply upload subtitle and video files to automatically match text to video content. Upload video and raw transcript file without timeline information. Timestamps will be automatically added to the transcriptions. Edit it online, then download subtitle files or video with subtitles directly. Original video language supports English, Spanish, Simplified Chinese, Traditional Chinese, Cantonese, Japanese, Korean, French, Thai, Russian, Portuguese, German, Italian, Vietnamese, Arabic. Single line word limit means the maximum number of words in a line of subtitles. When a paragraph contains many words, the system will make reasonable cuts according to the single line word limit to ensure that the number of words in a line of subtitles does not exceed the limit, therefore improving the subtitle display and facilitating reading.
Learn more
Alibaba Cloud Intelligent Speech Interaction
Intelligent Speech Interaction is developed based on state-of-the-art technologies such as speech recognition, speech synthesis, and natural language understanding. Enterprises can integrate Intelligent Speech Interaction into their products to enable them to listen, understand, and converse with users, providing users with an immersive human-computer interaction experience. Intelligent Speech Interaction is currently available in Mandarin Chinese, Cantonese Chinese, English, Japanese, Korean, French and Indonesian, and please stay tuned for other languages. Intelligent Speech Interaction is suitable for various scenarios, including intelligent Q&A, intelligent quality inspection, real-time subtitling for speeches, and transcription of audio recordings. Intelligent Speech Interaction has been successfully applied in many industries such as finance, insurance, eCommerce and smart home.
Learn more
Voqusa
Voqusa is a free AI transcript generator that turns any video into accurate text for TikTok, YouTube, Instagram, Facebook, X, LinkedIn, and Pinterest. Users can paste a video link or upload audio or video, then get a clean transcript in seconds. Voqusa’s AI extracts speech, applies punctuation, and produces a readable transcript that can be copied, downloaded, translated into 14+ languages, or used directly in a content workflow. It supports 7 social platforms, YouTube long-form, and 80+ source languages, including English, Spanish, Japanese, Korean, Arabic, Mandarin, and Traditional Chinese, with automatic language detection and no language picker required. It runs entirely in the browser, with no extension, app, or software installation required. It helps creators and marketers analyze viral content patterns, build competitor swipe files, repurpose video content across platforms, turn videos into blog posts, captions, scripts, and threads, and search competitor transcripts.
Learn more