Robust Speech Recognition via Large-Scale Weak Supervision
Multilingual speech recognition and audio understanding model
Uses Qwen3-ASR, local LLM, Whisper, TEN-VAD
Fast multimodal LLM for real-time voice interaction and AI apps
The no-nonsense RAG chunking library
Open source AI VTuber platform with voice chat and Live2D avatars
Research and application of technologies such as nl processing
The media player for language learning, with dual subtitles
Contexts Optical Compression
Large Audio Language Model built for natural interactions
LLM Large Model of Selling Anchor
AI-powered tool for generating, optimizing, and translating subtitles
A proof-of-concept jupyter extension which converts english queries
Advanced NLP with spaCy: A free online course
Flock is a workflow-based low-code platform for building chatbots
Framework for building real-time voice and multimodal AI agents
Run local LLMs like llama, deepseek, kokoro etc. inside your browser
Production ready toolkit to run AI locally
Based on the LangChain/LangGraph framework
Qwen3-ASR is an open-source series of ASR models
Chinese XLNet pre-trained model
Real-time voice interactive digital human
GLM-4.6V/4.5V/4.1V-Thinking, towards versatile multimodal reasoning
Multilingual Document Layout Parsing in a Single Vision-Language Model
An on-premises, OCR-free unstructured data extraction