LTX-Video Support for ComfyUI
Unified Multimodal Understanding and Generation Models
Reference PyTorch implementation and models for DINOv3
CogView4, CogView3-Plus and CogView3(ECCV 2024)
Generating Immersive, Explorable, and Interactive 3D Worlds
Python inference and LoRA trainer package for the LTX-2 audio–video
Official implementation of Watermark Anything with Localized Messages
GLM-4.6V/4.5V/4.1V-Thinking, towards versatile multimodal reasoning
VGGSfM: Visual Geometry Grounded Deep Structure From Motion
State-of-the-art Image & Video CLIP, Multimodal Large Language Models
ICLR2024 Spotlight: curation/training code, metadata, distribution
Large-language-model & vision-language-model based on Linear Attention
PyTorch implementation of MAE