Tencent Hunyuan Multimodal diffusion transformer (MM-DiT) model
Personalize Any Characters with a Scalable Diffusion Transformer
Foundational video generation model with 13.6B parameters
Handwritten Text Recognition (HTR) system implemented with TensorFlow
OCRmyPDF adds an OCR text layer to scanned PDF files
Awesome multilingual OCR toolkits based on PaddlePaddle
Contexts Optical Compression
OCR software, free and offline
AI Agent Application Development Framework
A framework to enable multimodal models to operate a computer
SOTA Open Source TTS
Open source AI VTuber platform with voice chat and Live2D avatars
Industrial-level controllable zero-shot text-to-speech system
Automated translation solution for visual novels
Visual Causal Flow
Official inference repo for FLUX.2 models
Infinite Worlds with Versatile Interactions
HY-Motion model for 3D character animation generation
An on-premises, OCR-free unstructured data extraction
Accurate × Fast × Comprehensive
A Multi-Modal World Model for Reconstructing, Generating, Simulation
A Powerful Native Multimodal Model for Image Generation
MOSS‑TTS Family open‑source speech and sound generation model
A ranked list of awesome machine learning Python libraries
Converts text to speech in realtime