Accurate × Fast × Comprehensive
OCR expert VLM powered by Hunyuan's native multimodal architecture
A GPT-4o Level MLLM for Vision, Speech and Multimodal Live Streaming
Layout-aware OCR model for multilingual document understanding
Lightweight multimodal translation model for 55 languages