A Family of Open Sourced Music Foundation Models
Accurate × Fast × Comprehensive
OCR expert VLM powered by Hunyuan's native multimodal architecture
A GPT-4o Level MLLM for Vision, Speech and Multimodal Live Streaming
Lightweight multimodal translation model for 55 languages
Multimodal Transformer for document image understanding and layout
Layout-aware OCR model for multilingual document understanding
Qwen2.5-VL-3B-Instruct: Multimodal model for chat, vision & video
ClinicalBERT model trained on MIMIC notes for clinical NLP tasks
Small 3B-base multimodal model ideal for custom AI on edge hardware