State-of-the-art Image & Video CLIP, Multimodal Large Language Models
The world's only naturally intelligent knowledge technology
Learning multi-scale deep model correcting over- and under- exposed
Let us control diffusion models
Navigation mesh generation and pathfinding toolkit for game AI systems
Implementation of BEVFormer, a camera-only framework
Code release for "Masked-attention Mask Transformer
The official pytorch implementation of our paper
Pytorch framework for doing deep learning on point clouds
Estimates the psychovisual difference between two images
A Neural Net Training Interface on TensorFlow, with focus on speed
A PyTorch implementation of the NIPS 2017 paper
Code for "Image Generation from Scene Graphs", Johnson et al, CVPR 201
Style transfer, deep learning, feature transform
Code that accompanies my blog post outlining five video classification
Cross Audio-Visual Recognition using 3D Architectures
SPatial Analysis With self-organizing Neural Networks
Bayesian Program Learning model for one-shot learning
Software tool for designing spatial bar structures
Signal Processing and Classification Environment in Python using YAML
Multimodal Transformer for document image understanding and layout