Agentic, Reasoning, and Coding (ARC) foundation models
Long-form streaming TTS system for multi-speaker dialogue generation
MOSS‑TTS Family open‑source speech and sound generation model
Codex plugin that turns attached object images into code-only
Clean and efficient FP8 GEMM kernels with fine-grained scaling
OCR expert VLM powered by Hunyuan's native multimodal architecture
Production-tested AI infrastructure tools
Large Multimodal Models for Video Understanding and Editing
GLM-4.6V/4.5V/4.1V-Thinking, towards versatile multimodal reasoning
Open Multilingual Multimodal Chat LMs
Code for "Image Generation from Scene Graphs", Johnson et al, CVPR 201