A series of math-specific large language models of our Qwen2 series
The Clay Foundation Model - An open source AI model and interface
MOSS‑TTS Family open‑source speech and sound generation model
Bidirectional token-classification model for identifiable info
General-purpose image editing model that delivers high-fidelity
A Production-ready Reinforcement Learning AI Agent Library
A PyTorch library for implementing flow matching algorithms
PyTorch code and models for the DINOv2 self-supervised learning
Fast-stable-diffusion + DreamBooth
Collection of Gemma 3 variants that are trained for performance
VMZ: Model Zoo for Video Modeling
Official implementation of Watermark Anything with Localized Messages
Multimodal Diffusion with Representation Alignment
A Multi-Modal World Model for Reconstructing, Generating, Simulation
Controllable & emotion-expressive zero-shot TTS
A Systematic Framework for Interactive World Modeling
Unified Multimodal Understanding and Generation Models
DeepMind model for tracking arbitrary points across videos & robotics
Mixture-of-Experts Vision-Language Models for Advanced Multimodal
A Powerful Native Multimodal Model for Image Generation
Generating Immersive, Explorable, and Interactive 3D Worlds
Use ChatGPT to summarize the arXiv papers
OCR expert VLM powered by Hunyuan's native multimodal architecture
Project Lyra: Open Generative 3D World Models
Ultra-Efficient LLMs on End Device