PyTorch code and models for V-JEPA self-supervised learning from video
Official code base for LeWorldModel: Stable End-to-End Joint-Embedding
Mixture-of-Experts Vision-Language Models for Advanced Multimodal
End-to-end pipeline converting generative videos
CoTracker is a model for tracking any point (pixel) on a video
A Strong and Easy-to-use Single View 3D Hand+Body Pose Estimator
Joint Face Detection and Alignment
WaveRNN Vocoder + TTS
The official pytorch implementation of our paper
A real-time approach for mapping all human pixels of 2D RGB images
Efficient 3D human pose estimation in video using 2D keypoint
Deep Hough Voting for 3D Object Detection in Point Clouds