PyTorch code and models for V-JEPA self-supervised learning from video
Official code base for LeWorldModel: Stable End-to-End Joint-Embedding
Mixture-of-Experts Vision-Language Models for Advanced Multimodal
End-to-end pipeline converting generative videos
Serving system for machine learning models
A tool for learning vector representations of words and entities
CoTracker is a model for tracking any point (pixel) on a video
Meta-Transformer for Unified Multimodal Learning
Hypergraph Transformer for Skeleton-based Action Recognition
A Strong and Easy-to-use Single View 3D Hand+Body Pose Estimator
Joint Face Detection and Alignment
WaveRNN Vocoder + TTS
The official pytorch implementation of our paper
A real-time approach for mapping all human pixels of 2D RGB images
Graph embedding, classification and representation learning papers
An implementation of Tacotron 2 that supports multilingual experiments
Efficient 3D human pose estimation in video using 2D keypoint
Deep Hough Voting for 3D Object Detection in Point Clouds