Visual Causal Flow
Advancing Open-source World Models
Diversity-driven optimization and large-model reasoning ability
Official code base for LeWorldModel: Stable End-to-End Joint-Embedding
An experimental version of DeepSeek model
Infinite Worlds with Versatile Interactions
A SOTA open-source image editing model
Codex plugin that turns attached object images into code-only
High-Fidelity and Controllable Generation of Textured 3D Assets
Ling is a MoE LLM provided and open-sourced by InclusionAI
CLIP, Predict the most relevant text snippet given an image
code for Mesh R-CNN, ICCV 2019
Open-source image generative foundation model
Recovering the Visual Space from Any Views
4M: Massively Multimodal Masked Modeling
Accurate × Fast × Comprehensive
OCR expert VLM powered by Hunyuan's native multimodal architecture
Designed for text embedding and ranking tasks
Repo for SeedVR2 & SeedVR
Large Multimodal Models for Video Understanding and Editing
Implementation of the Surya Foundation Model for Heliophysics
Collection of Gemma 3 variants that are trained for performance
Pretrained time-series foundation model developed by Google Research
tiktoken is a fast BPE tokeniser for use with OpenAI's models
Repo of Qwen2-Audio chat & pretrained large audio language model