Framework for building real-time voice and multimodal AI agents
A full spaCy pipeline and models for scientific/biomedical documents
Image polygonal annotation with Python
Image processing in Python
Powerful Android AI agent with tools, automation, and Linux shell
Open source AI VTuber platform with voice chat and Live2D avatars
AsrTools: Smart Voice-to-Text Tool
Open Source Computer Vision Library
Voice Recognition to Text Tool
Fast multimodal LLM for real-time voice interaction and AI apps
Multilingual Automatic Speech Recognition with word-level timestamps
A high-quality tool for convert PDF to Markdown and JSON
From Addition, Subtraction, Multiplication, and Division to ML
LLM Large Model of Selling Anchor
OCRmyPDF adds an OCR text layer to scanned PDF files
Open source annotation tool for machine learning practitioners
Enhances Tesseract OCR output using LLMs (local or API)
OCR expert VLM powered by Hunyuan's native multimodal architecture
Omnilingual ASR Open-Source Multilingual SpeechRecognition
Library for OCR-related tasks powered by Deep Learning
Toolkit for conversational AI
Recognition and resolution of numbers, units, date/time, etc.
Han Language Processing
VoiceStudio is the open-source, fully-local ElevenLabs alternative
2D and 3D Face alignment library build using pytorch