Qwen3-omni is a natively end-to-end, omni-modal LLM
Repo of Qwen2-Audio chat & pretrained large audio language model
OpenTinker is an RL-as-a-Service infrastructure for foundation models
Tongyi Deep Research, the Leading Open-source Deep Research Agent
This repository contains the official implementation of FastVLM
CogView4, CogView3-Plus and CogView3(ECCV 2024)
Research code artifacts for Code World Model (CWM)
Chinese and English multimodal conversational language model
Reproduction of Poetiq's record-breaking submission to the ARC-AGI-1
Qwen2.5-VL is the multimodal large language model series
MedicalGPT: Training Your Own Medical GPT Model with ChatGPT Training
Diversity-driven optimization and large-model reasoning ability
Open-source multi-speaker long-form text-to-speech model
Open-weight, large-scale hybrid-attention reasoning model
Capable of understanding text, audio, vision, video
Open-source industrial-grade ASR models
Ling-V2 is a MoE LLM provided and open-sourced by InclusionAI
NVIDIA Isaac GR00T N1.5 is the world's first open foundation model
GUI shell for running local LLM on desktop
OCR expert VLM powered by Hunyuan's native multimodal architecture
LLM-based Reinforcement Learning audio edit model
CodeGeeX: An Open Multilingual Code Generation Model (KDD 2023)
CodeGeeX2: A More Powerful Multilingual Code Generation Model
ChatGLM-6B: An Open Bilingual Dialogue Language Model
Open-source, high-performance Mixture-of-Experts large language model