Visual intelligence for your home.
Open Vision Agents by Stream. Build voice and vision agents quickly
Enable AI to control your desktop, mobile and HMI devices
Implementation of Vision Transformer, a simple way to achieve SOTA
Open source framework for deep learning satellite and aerial imagery
Open Source Differentiable Computer Vision Library
Build Vision Agents quickly with any model or video provider
Phi-3.5 for Mac: Locally-run Vision and Language Models
Effortless data labeling with AI support from Segment Anything
3D reconstruction software
Witness the aha moment of VLM with less than $3
The repository provides code for running inference with SAM 2
Datasets, transforms and models specific to Computer Vision
Suite of reference architectures for building GPU-accelerated vision
Mixture-of-Experts Vision-Language Models for Advanced Multimodal
A lightweight vision library for performing large object detection
Fast image augmentation library and an easy-to-use wrapper
Automatically find issues in image datasets
Skywork-R1V is an advanced multimodal AI model series
State-of-the-art Image & Video CLIP, Multimodal Large Language Models
Unified web UI for training and running open models locally
An open phone agent model & framework
Advanced AI Explainability for computer vision
ICLR2024 Spotlight: curation/training code, metadata, distribution
Medical imaging toolkit for deep learning