Qwen-Image is a powerful image generation foundation model
MiniSom is a minimalistic implementation of the Self Organizing Maps
Sharp Monocular Metric Depth in Less Than a Second
AI tool converting video/audio into structured documents instantly
Programmatic access to the AlphaGenome model
Reverse engineering Gemini's SynthID detection
Aider is AI pair programming in your terminal
PyTorch extensions for fast R&D prototyping and Kaggle farming
Recovering the Visual Space from Any Views
Opensource browser using agents
Unofficial Python API and agentic skill for Google NotebookLM
Simple and easily configurable grid world environments
A Unified Framework for Text-to-3D and Image-to-3D Generation
A Python package for segmenting geospatial data with the SAM
Synthesizing and manipulating 2048x1024 images with conditional GANs
MapAnything: Universal Feed-Forward Metric 3D Reconstruction
Multimodal embedding and reranking models built on Qwen3-VL
Tool for exploring and debugging transformer model behaviors
A tension reasoning engine over 131 S-class problems
Fast3R: Towards 3D Reconstruction of 1000+ Images in One Forward Pass
Tooling for the Common Objects In 3D dataset
VGGSfM: Visual Geometry Grounded Deep Structure From Motion
Generate 3D objects conditioned on text or images
Let us control diffusion models
Meta-Transformer for Unified Multimodal Learning