GLM-4 series: Open Multilingual Multimodal Chat LMs
Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence
Official code base for LeWorldModel: Stable End-to-End Joint-Embedding
Tiny vision language model
Mixture-of-Experts Vision-Language Models for Advanced Multimodal
scikit-learn compatible tabular foundation model
Model export recipes, Python primitives, and Swift runtime utilities
PyTorch code and models for the DINOv2 self-supervised learning
A 0.1B Omni model trained from scratch
26m function call model that runs on incredibly small devices
Video Object and Interaction Deletion
Video understanding codebase from FAIR for reproducing video models
Audio foundation model excelling in audio understanding
A Systematic Framework for Interactive World Modeling
DeepSeek Coder: Let the Code Write Itself
1B text generation model based on the HRM architecture
Robust Speech Recognition Across Languages, Dialects
Open-Source Financial Large Language Models
Repo for SeedVR2 & SeedVR
A Customizable Image-to-Video Model based on HunyuanVideo
Repo of Qwen2-Audio chat & pretrained large audio language model
Tongyi Deep Research, the Leading Open-source Deep Research Agent
Qwen2.5-VL is the multimodal large language model series
Generate Any 3D Scene in Seconds
Advancing Open-source World Models