Python inference and LoRA trainer package for the LTX-2 audio–video
From Addition, Subtraction, Multiplication, and Division to ML
SAPIEN Manipulation Skill Framework
Driving with Graph Visual Question Answering
LISA: Reasoning Segmentation via Large Language Model
This repository contains the official implementation of FastVLM
Refer and Ground Anything Anywhere at Any Granularity
Self-supervised visual learning using momentum contrast in PyTorch
Contexts Optical Compression
Misc; latest version of waifu2x; 2D video to stereo 3D video
Wan2.1: Open and Advanced Large-Scale Video Generative Model
PS2 Covers Collection
Powerful framework for controlling Android and iOS devices
Make any agent harness multimodal-native
Video Object and Interaction Deletion
Master the fundamentals of machine learning, deep learning
Recovering the Visual Space from Any Views
AI-Powered Wiki Generator for GitHub/Gitlab/Bitbucket Repositories
VMZ: Model Zoo for Video Modeling
Taming Stable Diffusion for Lip Sync
Foundational video generation model with 13.6B parameters
AI tool that converts GitHub repositories into interactive diagrams
CogView4, CogView3-Plus and CogView3(ECCV 2024)
[CVPR 2026 Oral] VGGT Omega
Reference PyTorch implementation and models for DINOv3