Visual Causal Flow
Designed for text embedding and ranking tasks
Diversity-driven optimization and large-model reasoning ability
Recovering the Visual Space from Any Views
PyTorch code and models for the DINOv2 self-supervised learning
Block Diffusion for Ultra-Fast Speculative Decoding
LTX-Video Support for ComfyUI
Infinite Worlds with Versatile Interactions
A SOTA open-source image editing model
High-Fidelity and Controllable Generation of Textured 3D Assets
Implementation of the Surya Foundation Model for Heliophysics
Codex plugin that turns attached object images into code-only
Ling is a MoE LLM provided and open-sourced by InclusionAI
Collection of Gemma 3 variants that are trained for performance
CLIP, Predict the most relevant text snippet given an image
GLM-4.5: Open-source LLM for intelligent agents by Z.ai
Advancing Open-source World Models
Repo for SeedVR2 & SeedVR
Pretrained time-series foundation model developed by Google Research
Open-source image generative foundation model
4M: Massively Multimodal Masked Modeling
Global weather forecasting model using graph neural networks and JAX
OCR expert VLM powered by Hunyuan's native multimodal architecture
tiktoken is a fast BPE tokeniser for use with OpenAI's models
Large Multimodal Models for Video Understanding and Editing