Robust Speech Recognition via Large-Scale Weak Supervision
Speech-to-text, text-to-speech, and speaker recognition
Multilingual speech recognition and audio understanding model
OpenVINO™ Toolkit repository
Omnilingual ASR Open-Source Multilingual SpeechRecognition
Cross-platform AI language practice app
Han Language Processing
Replace OpenAI GPT with another LLM in your app
Underthesea - Vietnamese NLP Toolkit
Build your own AI friend
A full spaCy pipeline and models for scientific/biomedical documents
Uses Qwen3-ASR, local LLM, Whisper, TEN-VAD
Enhances Tesseract OCR output using LLMs (local or API)
A GUI Agent app based on UI-TARS to control your computer using AI
Toolkit for conversational AI
Fast multimodal LLM for real-time voice interaction and AI apps
The no-nonsense RAG chunking library
Open source AI VTuber platform with voice chat and Live2D avatars
A PyTorch-based Speech Toolkit
Training data (data labeling, annotation, workflow) for all data types
A pure Javascript Multilingual OCR
Research and application of technologies such as nl processing
The media player for language learning, with dual subtitles
Recognition and resolution of numbers, units, date/time, etc.
Contexts Optical Compression