Robust Speech Recognition via Large-Scale Weak Supervision
Multilingual speech recognition and audio understanding model
Omnilingual ASR Open-Source Multilingual SpeechRecognition
Han Language Processing
Replace OpenAI GPT with another LLM in your app
Underthesea - Vietnamese NLP Toolkit
A full spaCy pipeline and models for scientific/biomedical documents
Uses Qwen3-ASR, local LLM, Whisper, TEN-VAD
Enhances Tesseract OCR output using LLMs (local or API)
Toolkit for conversational AI
Fast multimodal LLM for real-time voice interaction and AI apps
The no-nonsense RAG chunking library
Open source AI VTuber platform with voice chat and Live2D avatars
A PyTorch-based Speech Toolkit
Training data (data labeling, annotation, workflow) for all data types
Contexts Optical Compression
Repo of Qwen2-Audio chat & pretrained large audio language model
kaldi-asr/kaldi is the official location of the Kaldi project
Library for OCR-related tasks powered by Deep Learning
Large Audio Language Model built for natural interactions
Persian NLP Toolkit
LLM Large Model of Selling Anchor
Integrating LLMs into structured NLP pipelines
Semantic search and workflows for medical/scientific papers
AI-powered tool for generating, optimizing, and translating subtitles