Official code base for LeWorldModel: Stable End-to-End Joint-Embedding
Visual Causal Flow
Tiny vision language model
GLM-4.6V/4.5V/4.1V-Thinking, towards versatile multimodal reasoning
Open-source image generative foundation model
Video Object and Interaction Deletion
Qwen3-ASR is an open-source series of ASR models
Foundation model for image generation
Fast-stable-diffusion + DreamBooth
Z80-μLM is a 2-bit quantized language model
4M: Massively Multimodal Masked Modeling
Official implementation of DreamCraft3D
New family of code large language models (LLMs)
A Systematic Framework for Interactive World Modeling
FAIR Sequence Modeling Toolkit 2
Renderer for the harmony response format to be used with gpt-oss
Repo of Qwen2-Audio chat & pretrained large audio language model
Tongyi Deep Research, the Leading Open-source Deep Research Agent
Phi-3.5 for Mac: Locally-run Vision and Language Models
Stable Diffusion WebUI Forge is a platform on top of Stable Diffusion
GLM-4.6V/4.5V/4.1V-Thinking, towards versatile multimodal reasoning
MedicalGPT: Training Your Own Medical GPT Model with ChatGPT Training
Netease Youdao's open-source embedding and reranker models
An Efficient Agentic Model for Computer Use
Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence