High-Resolution Image Synthesis with Latent Diffusion Models
High-Resolution 3D Assets Generation with Large Scale Diffusion Models
LLM-based Reinforcement Learning audio edit model
FAIR Sequence Modeling Toolkit 2
MOSS‑TTS Family open‑source speech and sound generation model
Diffusion Transformer with Fine-Grained Chinese Understanding
GLM-4.6V/4.5V/4.1V-Thinking, towards versatile multimodal reasoning
CogView4, CogView3-Plus and CogView3(ECCV 2024)
A Systematic Framework for Interactive World Modeling
The official repo of Qwen chat & pretrained large language model
Open-source framework for intelligent speech interaction
OCR expert VLM powered by Hunyuan's native multimodal architecture
Unified Multimodal Understanding and Generation Models
Repo of Qwen2-Audio chat & pretrained large audio language model
Qwen3-ASR is an open-source series of ASR models
Implementation of "MobileCLIP" CVPR 2024
Qwen3 is the large language model series developed by Qwen team
Block Diffusion for Ultra-Fast Speculative Decoding
Bidirectional token-classification model for identifiable info
Ultra-Efficient LLMs on End Device
Generate Any 3D Scene in Seconds
A GPT-4o Level MLLM for Vision, Speech and Multimodal Live Streaming
Chinese and English multimodal conversational language model
Audio foundation model excelling in audio understanding
Multi-modal large language model designed for audio understanding