Mixture-of-Experts Vision-Language Models for Advanced Multimodal
Long-form streaming TTS system for multi-speaker dialogue generation
General-purpose image editing model that delivers high-fidelity
PyTorch code and models for the DINOv2 self-supervised learning
One-click local MCP server installation in desktop apps
A Customizable Image-to-Video Model based on HunyuanVideo
Open-source large language model family from Tencent Hunyuan
Personalize Any Characters with a Scalable Diffusion Transformer
C++ implementation of ChatGLM-6B & ChatGLM2-6B & ChatGLM3 & GLM4(V)
Phi-3.5 for Mac: Locally-run Vision and Language Models
INT4/INT5/INT8 and FP16 inference on CPU for RWKV language model
Qwen-Image-Layered: Layered Decomposition for Inherent Editablity
GLM-4-Voice | End-to-End Chinese-English Conversational Model
1B text generation model based on the HRM architecture
Official code base for LeWorldModel: Stable End-to-End Joint-Embedding
Accurate × Fast × Comprehensive
Repo of Qwen2-Audio chat & pretrained large audio language model
Qwen2.5-VL is the multimodal large language model series
Advancing Open-source World Models
Open image model at the forefront of design
26m function call model that runs on incredibly small devices
Qwen3-ASR is an open-source series of ASR models
Collection of Gemma 3 variants that are trained for performance
Tool for exploring and debugging transformer model behaviors
CLIP, Predict the most relevant text snippet given an image