A SOTA open-source image editing model
Unified Multimodal Understanding and Generation Models
Sharp Monocular Metric Depth in Less Than a Second
Genome modeling and design across all domains of life
Achieving 3+ generation speedup on reasoning tasks
Ultra-Efficient LLMs on End Device
Official code base for LeWorldModel: Stable End-to-End Joint-Embedding
Open-Source Financial Large Language Models
Open-source large language model family from Tencent Hunyuan
Tongyi Deep Research, the Leading Open-source Deep Research Agent
State-of-the-art Image & Video CLIP, Multimodal Large Language Models
GLM-4-Voice | End-to-End Chinese-English Conversational Model
GPT4V-level open-source multi-modal model based on Llama3-8B
GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning
Renderer for the harmony response format to be used with gpt-oss
Open image model at the forefront of design
Foundation model for image generation
A Pragmatic VLA Foundation Model
VMZ: Model Zoo for Video Modeling
Video understanding codebase from FAIR for reproducing video models
Tool for exploring and debugging transformer model behaviors
Bidirectional token-classification model for identifiable info
Project Lyra: Open Generative 3D World Models
Inference script for Oasis 500M
Fast and Universal 3D reconstruction model for versatile tasks