Refer and Ground Anything Anywhere at Any Granularity
PyTorch code and models for VJEPA2 self-supervised learning from video
Language modeling in a sentence representation space
Bailing is a voice dialogue robot similar to GPT-4o
GLM-4.6V/4.5V/4.1V-Thinking, towards versatile multimodal reasoning
fast C++ library for linear algebra & scientific computing
Air traffic control tower and radar simulator (solo + multi-player)
Cerberus Content Management System
Jupyter notebooks from the scikit-learn video series
A fast, powerful, and simple hierarchical vision transformer
Resources, corpora, and tools for Chinese natural language processing
se GPT or other prompt based models to get structured output
Code release for ConvNeXt V2 model
Human Activity Recognition example using TensorFlow on smartphone
FaceXlib aims at providing ready-to-use face-related functions
Repository of notes, code and notebooks in Python
Code release for ConvNeXt model
A CLI script to generate subtitle files (SRT/VTT/TXT) for any video
Generate text images for training deep learning ocr model
A deep learning library for video understanding research
Fingerprint recognition gadget integrates multiple web libraries
The official pytorch implementation of our paper
Efficient 3D human pose estimation in video using 2D keypoint
A crash course in six episodes for software developers
Python TikTok bot