Showing 800 open source projects for "vision"

View related business solutions
  • Ship Agents Faster Icon
    Ship Agents Faster

    Transform your applications and workflows into powerful agentic systems at global scale.

    Gemini Enterprise Agent Platform lets you rapidly build, scale, govern and optimize production-ready agents grounded in your organization's data. The platform enables developers to build custom or pre-built agents for virtually any use case. New customers get $300 in free credits.
    Start Free
  • MongoDB Atlas runs apps anywhere Icon
    MongoDB Atlas runs apps anywhere

    Deploy in 115+ regions with the modern database for every enterprise.

    MongoDB Atlas gives you the freedom to build and run modern applications anywhere—across AWS, Azure, and Google Cloud. With global availability in over 115 regions, Atlas lets you deploy close to your users, meet compliance needs, and scale with confidence across any geography.
    Start Free
  • 1
    M2SQL

    M2SQL

    Replicate your Mumps Globals into Relational Data Base

    M2SQL is a handy tool that permits to map your existing Mumps Globals in a relational Vision. The mapping process is simple and intuitive and is done only once. The Mapping process will describe the Mumps Globals data as a complex of Applications, Hierarchies, Relations and fields. Once mapping is completed you will able to replicate all the Mumps data into Target Relation database such as SQLITE or MYSQL. Normally you will do the replication process once a day while system is silent. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 2
    DnCNN

    DnCNN

    Beyond a Gaussian Denoiser: Residual Learning of Deep CNN

    This repository implements DnCNN (“Deep CNN Denoiser”) from the paper “Beyond a Gaussian Denoiser: Residual Learning of Deep CNN for Image Denoising”. DnCNN is a feedforward convolutional neural network that learns to predict the residual noise (i.e. noise map) from a noisy input image, which is then subtracted to yield a clean image. This formulation allows efficient denoising, supports blind Gaussian noise (i.e. unknown noise levels), and can be extended to related tasks like image...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 3
    This project aims to be a simple and useful front-end for network rendering for POVRay (Persistence of Vision Raytracer) and in the next future for YafRay too.
    Downloads: 1 This Week
    Last Update:
    See Project
  • 4
    Machine Learning & Deep Learning

    Machine Learning & Deep Learning

    machine learning and deep learning tutorials, articles

    ...The project focuses on helping learners understand machine learning through hands-on coding examples rather than purely theoretical explanations. Each tutorial walks through the process of building and training models for tasks such as image classification, neural network training, and computer vision applications. The repository also includes explanations of how different algorithms function internally, helping readers connect theoretical knowledge with implementation details. Because the tutorials are organized into separate projects, users can easily explore specific topics or technologies within the machine learning ecosystem.
    Downloads: 0 This Week
    Last Update:
    See Project
  • $300 Free Credits to Build on Google Cloud Icon
    $300 Free Credits to Build on Google Cloud

    New customers can spin up VMs, build with AI, and query data at no cost.

    Put your $300 in credit toward real workloads, then keep building with free monthly usage for 20+ products. No commitment and no charge until you upgrade.
    Start Free
  • 5
    Deep Learning 500 Questions

    Deep Learning 500 Questions

    500 Questions on Deep Learning using a question-and-answer format

    ...The first sections focus on essential mathematics, machine learning basics, and deep learning foundations, establishing the groundwork for more advanced topics. Later chapters explore classic neural network structures such as CNNs, RNNs, and GANs, as well as key applications in computer vision like object detection and image segmentation. The resource also delves into optimization methods, including transfer learning, network architecture design, hyperparameter tuning, model compression, and acceleration techniques.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 6
    YOLOR

    YOLOR

    implementation of paper - You Only Learn One Representation

    ...YOLOR includes model configurations, training code, evaluation scripts, inference tools, and pretrained weights. Its central contribution is the use of implicit knowledge to improve network performance without treating every task as fully separate. It is useful for computer vision researchers and developers studying YOLO-style detectors, representation learning, and high-performance detection systems.
    Downloads: 1 This Week
    Last Update:
    See Project
  • 7
    qiji-font

    qiji-font

    Typeface from Ming Dynasty woodblock printed books

    Typeface from Ming Dynasty woodblock printed books. A Ming typeface. Extracted from Ming Dynasty woodblock printed books (凌閔刻本). Using semi-automatic computer vision and OCR. Open-source. A work in progress. Named in honor of 閔齊伋, a 16th-century printer. Intended to be used with Kenyan-lang, the Classical Chinese programming language. Download high-resolution PDFs and split pages into images. Manually lay a grid on top of each page to generate bounding boxes for characters (potentially replaceable by an automatic corner-detection algorithm). ...
    Downloads: 2 This Week
    Last Update:
    See Project
  • 8
    VRN

    VRN

    Code for "Large Pose 3D Face Reconstruction

    The VRN (Volumetric Regression Network) repository implements the “Large Pose 3D Face Reconstruction from a Single Image via Direct Volumetric CNN Regression” method. Instead of explicitly fitting a 3D model via landmark estimation and deformation, VRN treats the reconstruction task as volumetric segmentation: it learns a CNN to regress a 3D volume aligned to the input image, and then extracts a mesh via isosurface from that volume. The network is unguided (no 2D landmarks as intermediate)....
    Downloads: 0 This Week
    Last Update:
    See Project
  • 9
    CameraVoyeur

    CameraVoyeur

    Windows-oriented utility to log frames from a connected camera devices

    ...This makes it useful for developers who want to add camera capture to a Windows tool or who need a barebones sample to build surveillance, monitoring, or computer vision toys on top of. Its value is really in being a simple, readable reference rather than a big camera suite.
    Downloads: 0 This Week
    Last Update:
    See Project
  • Custom VMs From 1 to 96 vCPUs With 99.95% Uptime Icon
    Custom VMs From 1 to 96 vCPUs With 99.95% Uptime

    General-purpose, compute-optimized, or GPU/TPU-accelerated. Built to your exact specs.

    Live migration and automatic failover keep workloads online through maintenance. One free e2-micro VM every month.
    Start Free
  • 10
    CAM

    CAM

    Class Activation Mapping

    This repository implements Class Activation Mapping (CAM), a technique to expose the implicit attention of convolutional neural networks by generating heatmaps that highlight the most discriminative image regions influencing a network’s class prediction. The method involves modifying a CNN model slightly (e.g., using global average pooling before the final layer) to produce a weighted combination of feature maps as the class activation map. Integration with existing CNNs (with light...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 11
    Alabaster Theme

    Alabaster Theme

    A light theme for Visual Studio Code

    ...The repository documents the rationale in detail and has inspired ports to other editors, showing its appeal to developers who prefer quiet aesthetics. Community discussions in the issue tracker revolve around targeted tweaks rather than expanding the palette, consistent with the theme’s minimalist vision.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 12
    YOLOv4-large

    YOLOv4-large

    Scaled-YOLOv4: Scaling Cross Stage Partial Network

    YOLOv4-large is an open-source implementation of the Scaled-YOLOv4 object detection architecture, designed to improve both the accuracy and scalability of real-time computer vision models. The project provides a PyTorch implementation of the Scaled-YOLOv4 framework, which extends the original YOLOv4 architecture using Cross Stage Partial (CSP) networks and new scaling techniques. Unlike earlier object detection systems that only scale depth or width, this architecture scales multiple aspects of the neural network including structure, resolution, and channel configuration. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 13
    PyCls

    PyCls

    Codebase for Image Classification Research, written in PyTorch

    pycls is a focused PyTorch codebase for image classification research that emphasizes reproducibility and strong, transparent baselines. It popularized families like RegNet and supports classic architectures (ResNet, ResNeXt) with clean implementations and consistent training recipes. The repository includes highly tuned schedules, augmentations, and regularization settings that make it straightforward to match reported accuracy without guesswork. Distributed training and mixed precision are...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 14

    VIDGUI

    A simple, friendly framework for Free Pascal text mode GUI creation.

    A framework for creating Free Pascal text mode GUIs, it aims to be simple and friendly. Not as powerful as Free Vision but much easier to use.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 15
    TimeSformer

    TimeSformer

    The official pytorch implementation of our paper

    TimeSformer is a vision transformer architecture for video that extends the standard attention mechanism into spatiotemporal attention. The model alternates attention along spatial and temporal dimensions (or designs variants like divided attention) so that it can capture both appearance and motion cues in video. Because the attention is global across frames, TimeSformer can reason about dependencies across long time spans, not just local neighborhoods.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 16
    Nerfies

    Nerfies

    This is the code for Deformable Neural Radiance Fields

    ...A set of utilities manages dataset preparation, pose estimation, and checkpoints so researchers can reproduce results on their own footage. The work sits at the intersection of graphics and vision, showing how learned volumetric rendering can handle human motion without dense markers or studio rigs.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 17
    deep-learning-for-image-processing

    deep-learning-for-image-processing

    deep learning for image processing including classification

    deep-learning-for-image-processing is an extensive educational repository covering practical deep learning methods for computer vision. It organizes implementations, explanations, presentation files, and video lessons around major neural network architectures. The material teaches both model structure and training workflows, with examples built in PyTorch and TensorFlow through Keras. Classification topics range from LeNet and AlexNet to ResNet, EfficientNet, Vision Transformer, Swin Transformer, ConvNeXt, and MobileViT. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 18
    course-v3

    course-v3

    The 3rd edition of course.fast.ai

    course-v3 repository contains the complete learning materials for the third edition of the Practical Deep Learning for Coders course developed by the fast.ai research group. The repository includes Jupyter notebooks, lesson materials, datasets, and supporting documentation used in the course to teach modern deep learning techniques. The course emphasizes a top-down approach to learning artificial intelligence, where students begin by building practical models and later study the underlying...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 19
    YOLO ROS

    YOLO ROS

    YOLO ROS: Real-Time Object Detection for ROS

    This is a ROS package developed for object detection in camera images. You only look once (YOLO) is a state-of-the-art, real-time object detection system. In the following ROS package, you are able to use YOLO (V3) on GPU and CPU. The pre-trained model of the convolutional neural network is able to detect pre-trained classes including the data set from VOC and COCO, or you can also create a network with your own detection objects. The YOLO packages have been tested under ROS Noetic and...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 20
    DensePose

    DensePose

    A real-time approach for mapping all human pixels of 2D RGB images

    DensePose is a computer vision system that maps all human pixels in an RGB image to the 3D surface of a human body model. It extends human pose estimation from predicting joint keypoints to providing dense correspondences between 2D images and a canonical 3D mesh (such as the SMPL model). This enables detailed understanding of human shape, motion, and surface appearance directly from images or videos.
    Downloads: 3 This Week
    Last Update:
    See Project
  • 21
    Gluon CV Toolkit

    Gluon CV Toolkit

    Gluon CV Toolkit

    GluonCV provides implementations of state-of-the-art (SOTA) deep learning algorithms in computer vision. It aims to help engineers, researchers, and students quickly prototype products, validate new ideas and learn computer vision. It features training scripts that reproduce SOTA results reported in latest papers, a large set of pre-trained models, carefully designed APIs and easy-to-understand implementations and community support.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 22
    SLAMBook-en

    SLAMBook-en

    The English version of 14 lectures on visual SLAM

    This project is the English version of “14 Lectures on Visual SLAM: From Theory to Practice,” a text and teaching resource about visual simultaneous localization and mapping (SLAM). It provides the full LaTeX source (formerly Markdown) for all 14 chapters, letting readers compile and study the material systematically. Within the repository you’ll find organized subfolders (e.g. chapters, latex, resources) containing the lecture contents, references, figures, and supporting assets for each...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 23
    PyTorch SimCLR

    PyTorch SimCLR

    PyTorch implementation of SimCLR: A Simple Framework

    For quite some time now, we know about the benefits of transfer learning in Computer Vision (CV) applications. Nowadays, pre-trained Deep Convolution Neural Networks (DCNNs) are the first go-to pre-solutions to learn a new task. These large models are trained on huge supervised corpora, like the ImageNet. And most important, their features are known to adapt well to new problems. This is particularly interesting when annotated training data is scarce.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 24
    CNN for Image Retrieval
    ...It focuses on applying deep learning techniques to improve upon traditional handcrafted descriptors by learning features directly from data. The code includes training and evaluation scripts that can be adapted for custom datasets, making it useful for experimenting with retrieval systems in computer vision. By leveraging CNN architectures, the project showcases how learned embeddings can capture semantic similarity across varied images. This resource serves as both an educational reference and a foundation for further exploration in image retrieval research.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 25
    Minecraft

    Minecraft

    Simple Minecraft-inspired program using Python and Pyglet

    ...It provides full WASD movement with mouse look and basic building controls, giving a surprisingly complete “sandbox” feel in only a small amount of Python code. The author explicitly describes a vision for the project as an educational tool: because many kids love Minecraft and Python is a friendly first programming language, the codebase is meant to be readable, hackable, and heavily commented so learners can tweak parameters like gravity or walking speed and see the effects immediately. Cross-platform instructions are included, with guidance for macOS users who may need to run Python in 32-bit mode or install a specific Pyglet version for compatibility.
    Downloads: 100 This Week
    Last Update:
    See Project