Showing 800 open source projects for "vision"

View related business solutions
  • $300 Free Credits to Build on Google Cloud Icon
    $300 Free Credits to Build on Google Cloud

    New customers can spin up VMs, build with AI, and query data at no cost.

    Put your $300 in credit toward real workloads, then keep building with free monthly usage for 20+ products. No commitment and no charge until you upgrade.
    Start Free
  • Build Agents and Models on One Platform Icon
    Build Agents and Models on One Platform

    Everything you need to build production-ready agents and models. Access 200+ Google and third-party AI models and tools.

    Gemini Enterprise Agent Platform is Google Cloud's comprehensive platform for developers to build, scale, govern, and optimize agents and models. Choose from Google's most advanced models and third-party models like Anthropic's Claude Model Family.
    Start Free
  • 1
    Top deep learning Github repositories

    Top deep learning Github repositories

    Top 200 deep learning Github repositories sorted by stars

    ...Instead of providing its own machine learning models or frameworks, the project functions as an organized index that helps users discover high-quality deep learning repositories across different application domains. The repository categorizes projects related to neural networks, computer vision, natural language processing, reinforcement learning, and other areas of artificial intelligence. By collecting popular open-source implementations in one place, the project simplifies the process of exploring cutting-edge tools and research implementations for deep learning practitioners. The curated lists are particularly helpful for developers who want to quickly identify well-maintained projects with strong community support.
    Downloads: 2 This Week
    Last Update:
    See Project
  • 2

    Rapid Engineering Architecture Linked

    Architecture Template for Agile Development

    ...R-E-A-L (Rapid Engineering Architecture Linked) allows agile while being formal. R-E-A-L (Rapid Engineering Architecture Linked) is highly influenced by the NAFv3. * Mission&Vision statement for each artifact, stakeholder analysis (NAV) * capability based approach: UC used to express capability (NOV) * service oriented approach: UC implies a service (NSOV) R-E-A-L (Rapid Engineering Architecture Linked) utilizes Design by Contract as concept on various levels of abstraction.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 3
    Face-Detection-JavaScript

    Face-Detection-JavaScript

    Small browser-based computer vision demo

    Face-Detection-JavaScript is a small browser-based computer vision demo that uses JavaScript and face-api.js to analyze live webcam video. It loads face detection, facial landmark, face recognition, and facial expression models from a local models folder. Once the webcam stream starts, the app creates a canvas overlay and draws detection results directly over the video feed. It can identify face boxes, map facial landmark points, and display detected expressions in near real time. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 4
    ncvs

    ncvs

    Network Camera Vision System

    Fully featured IP camera management system built in Qt creator implementing FFMPEG on windows. Includes pre built binary libraries for easy compilation. Requires Qt, MS Visual Studio, NVIDIA Cuda drivers. There is an install program for the executable
    Downloads: 0 This Week
    Last Update:
    See Project
  • Custom VMs From 1 to 96 vCPUs With 99.95% Uptime Icon
    Custom VMs From 1 to 96 vCPUs With 99.95% Uptime

    General-purpose, compute-optimized, or GPU/TPU-accelerated. Built to your exact specs.

    Live migration and automatic failover keep workloads online through maintenance. One free e2-micro VM every month.
    Start Free
  • 5
    Deep Learning Drizzle

    Deep Learning Drizzle

    Drench yourself in Deep Learning, Reinforcement Learning

    Drench yourself in Deep Learning, Reinforcement Learning, Machine Learning, Computer Vision, and NLP by learning from these exciting lectures! Optimization courses which form the foundation for ML, DL, RL. Computer Vision courses which are DL & ML heavy. Speech recognition courses which are DL heavy. Structured Courses on Geometric, Graph Neural Networks. Section on Autonomous Vehicles. Section on Computer Graphics with ML/DL focus.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 6
    The Integrating Vision Toolkit (IVT) is a powerful and fast C++ computer vision library with an easy-to-use object-oriented architecture. It offers its own multi-platform GUI toolkit. OpenCV is integrated optionally. Website: http://ivt.sourceforge.net
    Downloads: 3 This Week
    Last Update:
    See Project
  • 7
    Deep Learning with PyTorch

    Deep Learning with PyTorch

    Latest techniques in deep learning and representation learning

    This course concerns the latest techniques in deep learning and representation learning, focusing on supervised and unsupervised deep learning, embedding methods, metric learning, convolutional and recurrent nets, with applications to computer vision, natural language understanding, and speech recognition. The prerequisites include DS-GA 1001 Intro to Data Science or a graduate-level machine learning course. To be able to follow the exercises, you are going to need a laptop with Miniconda (a minimal version of Anaconda) and several Python packages installed. The following instruction would work as is for Mac or Ubuntu Linux users, Windows users would need to install and work in the Git BASH terminal. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 8
    GAAS

    GAAS

    GAAS is an open-source program designed for fully autonomous VTOL

    GAAS, or Generalized Autonomy Aviation System, is an open-source autonomy framework for drones, VTOL aircraft, and experimental flying-car research. It focuses on robust autonomous navigation using lidar rather than relying primarily on vision. The system includes mapping, HD-map localization, path planning, obstacle handling, and aircraft-control modules. Localization components include NDT and ICP approaches, with GPU acceleration available for selected algorithms. Simulation support integrates lidar, stereo cameras, PX4, Gazebo, and AirSim environments. Its modular C++ architecture allows individual autonomy components to be replaced or extended. ...
    Downloads: 6 This Week
    Last Update:
    See Project
  • 9
    Downloads: 0 This Week
    Last Update:
    See Project
  • MongoDB Atlas runs apps anywhere Icon
    MongoDB Atlas runs apps anywhere

    Deploy in 115+ regions with the modern database for every enterprise.

    MongoDB Atlas gives you the freedom to build and run modern applications anywhere—across AWS, Azure, and Google Cloud. With global availability in over 115 regions, Atlas lets you deploy close to your users, meet compliance needs, and scale with confidence across any geography.
    Start Free
  • 10
    Miyo-Rolling
    From the creator of MiyoLinux, Miyo-Rolling is based on NuTyX GNU/Linux. While staying true to the vision of delivering a minimal and base system which allows the user to decide what goes on their system, Miyo-Rolling provides a rolling release that is extremely stable...thanks to the underlying NuTyX system. For now, it is only available as a 64bit system. If you would like to try the JWM version, it can be downloaded from the NuTyX GNU/Linux website...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 11
    GAAS

    GAAS

    Autonomous aviation intelligence software for drones and VTOL

    GAAS (Generalized Autonomy Aviation System) is an open source software platform for autonomous drones and VTOLs. GAAS was built to provide a common infrastructure for computer-vision based drone intelligence. In the long term, GAAS aims to accelerate the coming of autonomous VTOLs. Being a BSD-licensed product, GAAS makes it easy for enterprises, researches, and drone enthusiasts to modify the code to suit specific use cases. Our long-term vision is to implement GAAS in autonomous passenger carrying VTOLs (or "flying cars"). ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 12
    MMF

    MMF

    A modular framework for vision & language multimodal research

    MMF is a modular framework for vision and language multimodal research from Facebook AI Research. MMF contains reference implementations of state-of-the-art vision and language models and has powered multiple research projects at Facebook AI Research. MMF is designed from ground up to let you focus on what matters, your model, by providing boilerplate code for distributed training, common datasets and state-of-the-art pre-trained baselines out-of-the-box.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 13
    Activity Recognition

    Activity Recognition

    Resources about activity recognition

    ...The repository includes links to code in MATLAB, Python, summaries of algorithms, datasets, and relevant research papers. Feature extraction method summaries (e.g. motion, sensor, vision). Deep learning for activity recognition references.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 14

    OpenFace

    A state-of-the-art facial behavior analysis toolkit

    OpenFace is an advanced facial behavior analysis toolkit intended for computer vision and machine learning researchers, those in the affective computing community, and those who are simply interested in creating interactive applications based on facial behavior analysis. The OpenFace toolkit is capable of performing several complex facial analysis tasks, including facial landmark detection, eye-gaze estimation, head pose estimation and facial action unit recognition.
    Downloads: 10 This Week
    Last Update:
    See Project
  • 15
    NativeScript Documentation

    NativeScript Documentation

    Documentation, API reference, and code snippets for NativeScript

    NativeScript provides platform APIs directly to the JavaScript runtime (with strong types) for a rich TypeScript development experience. Building Web, iOS, Android, and Vision Pro apps with a shared codebase (aka, cross-platform apps) Building native platform apps with portable JavaScript skills. Augmenting JavaScript projects with platform API capabilities. AndroidTV and Watch development watchOS development. Learning native platforms through JavaScript understanding. Exploring platform API documentation by trying APIs directly from a web browser without requiring a platform development machine setup.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 16
    maskrcnn-benchmark

    maskrcnn-benchmark

    Fast, modular reference implementation of Instance Segmentation

    Mask R-CNN Benchmark is a PyTorch-based framework that provides high-performance implementations of object detection, instance segmentation, and keypoint detection models. Originally built to benchmark Mask R-CNN and related models, it offers a clean, modular design to train and evaluate detection systems efficiently on standard datasets like COCO. The framework integrates critical components—region proposal networks (RPNs), RoIAlign layers, mask heads, and backbone architectures such as...
    Downloads: 1 This Week
    Last Update:
    See Project
  • 17
    ChainerCV

    ChainerCV

    ChainerCV: a Library for Deep Learning in Computer Vision

    ChainerCV is a collection of tools to train and run neural networks for computer vision tasks using Chainer. In ChainerCV, we define the object detection task as a problem of, given an image, bounding box-based localization and categorization of objects. Bounding boxes in an image are represented as a two-dimensional array of shape (R,4), where R is the number of bounding boxes and the second axis corresponds to the coordinates of bounding boxes.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 18
    Replica Dataset

    Replica Dataset

    High-fidelity indoor 3D dataset for AI simulation and robotics

    Replica Dataset is a high-quality 3D dataset of realistic indoor environments designed to advance research in computer vision, robotics, and embodied AI. Developed by Facebook Research (now Meta AI), it features accurate geometric reconstructions, high-resolution and high dynamic range textures, and comprehensive semantic annotations. Each environment contains detailed models of real-world spaces, including rooms, furniture, glass, and mirror surfaces.
    Downloads: 12 This Week
    Last Update:
    See Project
  • 19
    GPUImage 2

    GPUImage 2

    Framework for GPU-accelerated video and image processing

    ...The original GPUImage framework was written in Objective-C and targeted Mac and iOS, but this latest version is written entirely in Swift and can also target Linux and future platforms that support Swift code. The objective of the framework is to make it as easy as possible to set up and perform realtime video processing or machine vision against image or video sources. By relying on the GPU to run these operations, performance improvements of 100X or more over CPU-bound code can be realized. This is particularly noticeable in mobile or embedded devices. On an iPhone 4S, this framework can easily process 1080p video at over 60 FPS. On a Raspberry Pi 3, it can perform Sobel edge detection on live 720p video at over 20 FPS.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 20
    JSON2YOLO

    JSON2YOLO

    Convert JSON annotations into YOLO format.

    ...Ultralytics is a U.S.-based particle physics and AI startup with over 6 years of expertise supporting government, academic, and business clients. We offer a wide range of vision AI services, spanning from simple expert advice up to the delivery of fully customized, end-to-end production solutions.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 21
    Software and HDL code for Elphel reconfigurable network cameras
    Downloads: 0 This Week
    Last Update:
    See Project
  • 22
    Butteraugli

    Butteraugli

    Estimates the psychovisual difference between two images

    butteraugli is a perceptual similarity metric designed to estimate how noticeable differences between two images will be to the human eye. Instead of simple pixel math, it models aspects of human vision—color sensitivity, spatial masking, and contrast perception—to highlight differences that viewers actually see. The core tool outputs a single “distance” score along with per-pixel or per-region maps that show where artifacts are most objectionable. These maps make it practical to tune compressor settings and confirm whether bitrate reductions are visually acceptable. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 23
    SSD

    SSD

    A PyTorch Implementation of Single Shot MultiBox Detector

    ...Its structure makes it useful both as a reference implementation for learning SSD and as a base for custom experimentation in detection research or practical computer vision projects.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 24
    ConvNet Burden

    ConvNet Burden

    Memory consumption and FLOP count estimates for convnets

    convnet-burden is a MATLAB toolbox / script collection estimating computational cost (FLOPs) and memory consumption of various convolutional neural network architectures. It lets users compute approximate burdens (in FLOPs, memory) for standard image classification CNN models (e.g. ResNet, VGG) based on network definitions. The tool helps researchers compare the computational efficiency of architectures or quantify resource needs. Estimation of memory consumption (e.g. feature map sizes,...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 25
    DetectAndTrack

    DetectAndTrack

    The implementation of an algorithm presented in the CVPR18 paper

    ...Although the repo has been archived and is now read-only, its issue tracker and artifacts remain useful for understanding implementation details and experimental settings. The project sits alongside other Facebook Research vision efforts, offering historical context for the evolution of video pose and tracking techniques. Researchers can still study the algorithms, adapt the pipeline, or port ideas into modern frameworks.
    Downloads: 0 This Week
    Last Update:
    See Project