Port of OpenAI's Whisper model in C/C++
State-of-the-art TTS model under 25MB
A GPT-4o Level MLLM for Vision, Speech and Multimodal Live Streaming
AI that sees your screen and listens to conversations
Transcribe and translate audio offline on your personal computer
Featuring powerful AI capabilities and supporting e-book formats
iOS application for Lumo
Reading book source
On-device Speech-to-Intent engine powered by deep learning
HTML5 js recording mp3 wav ogg webm amr format
Low-latency AI inference engine optimized for mobile devices
Offline speech recognition API for Android, iOS, Raspberry Pi
In-App assistant SDK to build a multimodal conversational UX for iOS
Supercharge your shortcuts
Easily apply cool gnarly voice filters to your audio files
Code examples for new APIs of iOS 10
Code examples for the new features of iOS 9
Easy AI Softwares for Blind, Deaf, Handicapped, Disabled People
Free & Easy AI Voice Accounting Software For Blind & Speechless People
State-of-the-art deep learning based audio codec
React Native Voice Recognition library for iOS and Android
SDK to build a multimodal conversational UX for Flutter apps
Build a multimodal conversational UX for apps created with React
Assistant SDK to build a multimodal conversational UX for Apache
Character animation system for games and simulations.