Convert AI papers to GUI
Large Multimodal Models for Video Understanding and Editing
OCR expert VLM powered by Hunyuan's native multimodal architecture
Tooling for the Common Objects In 3D dataset
A Customizable Image-to-Video Model based on HunyuanVideo
AI Suite for upscaling, interpolating & restoring images/videos
Implementation of a U-net complete with efficient attention
Implementation of Make-A-Video, new SOTA text to video generator
CoTracker is a model for tracking any point (pixel) on a video
A Strong and Easy-to-use Single View 3D Hand+Body Pose Estimator
Distributed Deep learning with Keras & Spark
CCTV Footage Timestamp Search Tool
Implementation of a Transformer based neural network
Pre-trained and Reproduced Deep Learning Models
The source code of CVPR 2019 paper "Deep Exemplar-based Colorization"
Reinforced Recommendation toolkit built around pytorch 1.7
The implementation of an algorithm presented in the CVPR18 paper
Code that accompanies my blog post outlining five video classification
Cross Audio-Visual Recognition using 3D Architectures