LTX-Video Support for ComfyUI
Visual Causal Flow
Code for running inference and finetuning with SAM 3 model
Official Python inference and LoRA trainer package
AI PPT Track Terminator, the strongest PPT Skill ever
Tiny vision language model
Open image model at the forefront of design
Qwen-Image-Layered: Layered Decomposition for Inherent Editablity
Unified Multimodal Understanding and Generation Models
Codex plugin that turns attached object images into code-only
Lets make video diffusion practical
Wan2.1: Open and Advanced Large-Scale Video Generative Model
Recovering the Visual Space from Any Views
This repository contains the official implementation of FastVLM
GLM-4.6V/4.5V/4.1V-Thinking, towards versatile multimodal reasoning
Python inference and LoRA trainer package for the LTX-2 audio–video
Video Object and Interaction Deletion
VMZ: Model Zoo for Video Modeling
Reference PyTorch implementation and models for DINOv3
Contexts Optical Compression
Mixture-of-Experts Vision-Language Models for Advanced Multimodal
GLM-4.6V/4.5V/4.1V-Thinking, towards versatile multimodal reasoning
Official implementation of Watermark Anything with Localized Messages
Multimodal Diffusion with Representation Alignment
Generating Immersive, Explorable, and Interactive 3D Worlds