Visual Causal Flow
LTX-Video Support for ComfyUI
Code for running inference and finetuning with SAM 3 model
Official Python inference and LoRA trainer package
AI PPT Track Terminator, the strongest PPT Skill ever
Tiny vision language model
Open image model at the forefront of design
Unified Multimodal Understanding and Generation Models
Codex plugin that turns attached object images into code-only
Lets make video diffusion practical
Qwen-Image-Layered: Layered Decomposition for Inherent Editablity
Python inference and LoRA trainer package for the LTX-2 audio–video
Contexts Optical Compression
This repository contains the official implementation of FastVLM
Wan2.1: Open and Advanced Large-Scale Video Generative Model
Video Object and Interaction Deletion
Recovering the Visual Space from Any Views
VMZ: Model Zoo for Video Modeling
CogView4, CogView3-Plus and CogView3(ECCV 2024)
Reference PyTorch implementation and models for DINOv3
Mixture-of-Experts Vision-Language Models for Advanced Multimodal
GLM-4.6V/4.5V/4.1V-Thinking, towards versatile multimodal reasoning
Official implementation of Watermark Anything with Localized Messages
Multimodal Diffusion with Representation Alignment
Generating Immersive, Explorable, and Interactive 3D Worlds