Visual intelligence for your home.
A refreshing functional take on deep learning
Official Repo For "Sa2VA: Marrying SAM2 with LLaVA
Pluggable SOTA multi-object tracking modules for segmentation
Gracefully face hCaptcha challenge with multimodal llms
Code for running inference with the SAM 3D Body Model 3DB
Python library and CLI tool to interface with Google Translate
A Python toolbox for scalable outlier detection
Create HTML profiling reports from pandas DataFrame objects
Superduper: Integrate AI models and machine learning workflows
A SOTA open-source image editing model
Qwen2.5-VL is the multimodal large language model series
Collections of robotics environments
Machine learning metrics for distributed, scalable PyTorch application
A self-hosted open source photo management service
A python module to repair invalid JSON from LLMs
VGGSfM: Visual Geometry Grounded Deep Structure From Motion
Marrying Grounding DINO with Segment Anything & Stable Diffusion
Simple and easily configurable grid world environments
Code release for Cut and Learn for Unsupervised Object Detection
Qwen-Image-Layered: Layered Decomposition for Inherent Editablity
Provides convenient access to the Anthropic REST API from any Python 3
JAX-based neural network library
Codex plugin that turns attached object images into code-only
Motion-controllable Video Generation via Latent Trajectory Guidance