Open-source image generative foundation model
Official code base for LeWorldModel: Stable End-to-End Joint-Embedding
Diversity-driven optimization and large-model reasoning ability
CLIP, Predict the most relevant text snippet given an image
A 0.1B Omni model trained from scratch
Large Multimodal Models for Video Understanding and Editing
Code for running inference and finetuning with SAM 3 model
Qwen3 is the large language model series developed by Qwen team
Infinite Worlds with Versatile Interactions
GLM-4 series: Open Multilingual Multimodal Chat LMs
A Customizable Image-to-Video Model based on HunyuanVideo
Accurate × Fast × Comprehensive
Visual Causal Flow
gpt-oss-120b and gpt-oss-20b are two open-weight language models
PyTorch code and models for the DINOv2 self-supervised learning
Code for running inference with the SAM 3D Body Model 3DB
Designed for text embedding and ranking tasks
Implementation of the Surya Foundation Model for Heliophysics
Codex plugin that turns attached object images into code-only
Ling is a MoE LLM provided and open-sourced by InclusionAI
An experimental version of DeepSeek model
Z80-μLM is a 2-bit quantized language model
Repo of Qwen2-Audio chat & pretrained large audio language model
High-Fidelity and Controllable Generation of Textured 3D Assets
4M: Massively Multimodal Masked Modeling