Turn WiFi signals into real-time human sensing and spatial awareness.
Spatial data processing for geomodeling
Medical imaging toolkit for deep learning
Refer and Ground Anything Anywhere at Any Granularity
Project Lyra: Open Generative 3D World Models
State-of-the-art (SoTA) text-to-video pre-trained model
Original reference implementation of "3D Gaussian Splatting"
A high performance implementation of HDBSCAN clustering
Official SeedVR2 Video Upscaler for ComfyUI
A Codex Skill for generating high-density, editable PowerPoints
Visual Causal Flow
Unsupervised Learning for Image Registration
3D plotting and mesh analysis through a streamlined interface
Models for object and human mesh reconstruction
Qwen-Image-Layered: Layered Decomposition for Inherent Editablity
Sharp Monocular View Synthesis in Less Than a Second
Video understanding codebase from FAIR for reproducing video models
A Systematic Framework for Interactive World Modeling
State-of-the-art Image & Video CLIP, Multimodal Large Language Models
Qwen2.5-VL is the multimodal large language model series
Unifying 3D Mesh Generation with Language Models
Gracefully face hCaptcha challenge with multimodal llms
code for Mesh R-CNN, ICCV 2019
SonicDive 8D Music Player v-1.0
An easy way to manage SQLite databases and query CSV files