The Iris Book: Addition, Subtraction, Multiplication, and Division
Robust Speech Recognition via Large-Scale Weak Supervision
Multilingual speech recognition and audio understanding model
Contexts Optical Compression
Audio foundation model excelling in audio understanding
Handwritten Text Recognition (HTR) system implemented with TensorFlow
High-Performance Face Recognition Library on PaddlePaddle & PyTorch
An Open-Source Toolkit for General-OCR Research and Applications
Uses Qwen3-ASR, local LLM, Whisper, TEN-VAD
Framework for building real-time voice and multimodal AI agents
Book_4_Matrix Power | The Iris Book: From Addition, Subtraction
Open source AI VTuber platform with voice chat and Live2D avatars
Fast multimodal LLM for real-time voice interaction and AI apps
From Addition, Subtraction, Multiplication, and Division to ML
LLM Large Model of Selling Anchor
Accurate × Fast × Comprehensive
A proof-of-concept jupyter extension which converts english queries
Detailed notes on HanLP's new book, "Introduction to Natural Language"
Video understanding codebase from FAIR for reproducing video models
The no-nonsense RAG chunking library
Open speech-to-speech models and pipelines by Hugging Face toolkit AI
Advanced NLP with spaCy: A free online course
Large Audio Language Model built for natural interactions
StreamSpeech is a seamless model for offline speech recognition
PRML algorithms implemented in Python