Personalize Any Characters with a Scalable Diffusion Transformer
Multimodal model achieving SOTA performance
A Multi-Modal World Model for Reconstructing, Generating, Simulation
Large-language-model & vision-language-model based on Linear Attention
Project Lyra: Open Generative 3D World Models
Unified Multimodal Understanding and Generation Models
PyTorch code and models for the DINOv2 self-supervised learning
A Systematic Framework for Interactive World Modeling
code for Mesh R-CNN, ICCV 2019
RGBD video generation model conditioned on camera input
Video understanding codebase from FAIR for reproducing video models
VGGSfM: Visual Geometry Grounded Deep Structure From Motion
AI-powered tool to quickly remove watermarks from images flawlessly
AI Suite for upscaling, interpolating & restoring images/videos
Detect faces in an image
Software that can generate photos from paintings
Code for the paper Hybrid Spectrogram and Waveform Source Separation
Code release for ConvNeXt V2 model
PyTorch implementation of MAE
Per-Pixel Classification is Not All You Need for Semantic Segmentation
A mix of GAN implementations including progressive growing
Efficient Image Captioning code in Torch, runs on GPU
LL model providing reasoning and conversational capabilities
CLIP ViT-bigG/14: Zero-shot image-text model trained on LAION-2B
Compact agentic model for coding, tools, and productivity tasks