Visual Causal Flow
Code for running inference and finetuning with SAM 3 model
LTX-Video Support for ComfyUI
Qwen-Image-Layered: Layered Decomposition for Inherent Editablity
Official Python inference and LoRA trainer package
Tiny vision language model
A state-of-the-art open visual language model
Recovering the Visual Space from Any Views
Unified Multimodal Understanding and Generation Models
Mixture-of-Experts Vision-Language Models for Advanced Multimodal
Python inference and LoRA trainer package for the LTX-2 audio–video
Official implementation of Watermark Anything with Localized Messages
This repository contains the official implementation of FastVLM
Contexts Optical Compression
GLM-4.6V/4.5V/4.1V-Thinking, towards versatile multimodal reasoning
Generating Immersive, Explorable, and Interactive 3D Worlds
Multimodal Diffusion with Representation Alignment
Wan2.1: Open and Advanced Large-Scale Video Generative Model
Lets make video diffusion practical
Reference PyTorch implementation and models for DINOv3
CogView4, CogView3-Plus and CogView3(ECCV 2024)
Foundation model for image generation
Video Object and Interaction Deletion
Phi-3.5 for Mac: Locally-run Vision and Language Models
GLM-4.6V/4.5V/4.1V-Thinking, towards versatile multimodal reasoning