GLM-4.6V/4.5V/4.1V-Thinking, towards versatile multimodal reasoning
C++ implementation of ChatGLM-6B & ChatGLM2-6B & ChatGLM3 & GLM4(V)
Audio Language Models are Few-Shot Learners
Open Source Speech Language Model
Open-source industrial-grade ASR models
Qwen3-ASR is an open-source series of ASR models
Foundation model for image generation
Block Diffusion for Ultra-Fast Speculative Decoding
Multimodal embedding and reranking models built on Qwen3-VL
Collection of Gemma 3 variants that are trained for performance
Implementation of "MobileCLIP" CVPR 2024
VMZ: Model Zoo for Video Modeling
Official implementation of Watermark Anything with Localized Messages
High-resolution models for human tasks
Tool for exploring and debugging transformer model behaviors
Ling is a MoE LLM provided and open-sourced by InclusionAI
A Unified Framework for Text-to-3D and Image-to-3D Generation
Multimodal Diffusion with Representation Alignment
Personalize Any Characters with a Scalable Diffusion Transformer
Run GLM-5.2 (744B MoE) on a 25GB-RAM consumer machine
Qwen3-omni is a natively end-to-end, omni-modal LLM
MOSS‑TTS Family open‑source speech and sound generation model
Bidirectional token-classification model for identifiable info
Genome modeling and design across all domains of life
Project Lyra: Open Generative 3D World Models