Showing 800 open source projects for "vision"

View related business solutions
  • MongoDB Atlas runs apps anywhere Icon
    MongoDB Atlas runs apps anywhere

    Deploy in 115+ regions with the modern database for every enterprise.

    MongoDB Atlas gives you the freedom to build and run modern applications anywhere—across AWS, Azure, and Google Cloud. With global availability in over 115 regions, Atlas lets you deploy close to your users, meet compliance needs, and scale with confidence across any geography.
    Start Free
  • Veeam Data Platform v13.1 - Get Your Free Trial Icon
    Veeam Data Platform v13.1 - Get Your Free Trial

    Secure by design, portable by default. Recover clean, fast, anywhere. Start a free trial.

    Try Veeam Data Platform today. Experience the unified platform that's secure by design, portable by default, and proven to recover clean, fast, and anywhere.
    Try it Free
  • 1
    daily-paper-computer-vision

    daily-paper-computer-vision

    Document papers compiled daily in computer vision/deep learning

    This repo is a running feed of computer-vision research, tracking new papers and notable results so practitioners can keep up without scouring multiple sites. It’s organized chronologically and often thematically, making it easy to scan what’s new in detection, segmentation, recognition, generative vision, 3D, and video understanding. The cadence is intentionally frequent, reflecting how quickly CV advances and how hard it is to maintain awareness while working full time. ...
    Downloads: 1 This Week
    Last Update:
    See Project
  • 2
    OpenFlamingo

    OpenFlamingo

    An open-source framework for training large multimodal models

    ...This repo is still under development, and we hope to release better-performing and larger OpenFlamingo models soon. If you have any questions, please feel free to open an issue. We also welcome contributions! We provide an initial OpenFlamingo 9B model using a CLIP ViT-Large vision encoder and a LLaMA-7B language model. In general, we support any CLIP vision encoder. For the language model, we support LLaMA, OPT, GPT-Neo, GPT-J, and Pythia models. OpenFlamingo is a multimodal language model that can be used for a variety of tasks. It is trained on a large multimodal dataset.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 3
    RoomGPT

    RoomGPT

    Upload a photo of your room to generate your dream room with AI

    RoomGPT is an open-source app that lets you upload a photo of your room and generate redesigned versions of it using AI. It uses a model such as ControlNet to condition the generation on the original room layout, producing realistic variations while preserving structure like walls, windows, and furniture placement. The app is built on Next.js and exposes a simple web interface where users can upload images, choose styles, and view generated outputs. Under the hood, it calls a hosted ML model...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 4
    SOD

    SOD

    An Embedded Computer Vision & Machine Learning Library

    SOD is an embedded, modern cross-platform computer vision and machine learning software library that expose a set of APIs for deep-learning, advanced media analysis & processing including real-time, multi-class object detection and model training on embedded systems with limited computational resource and IoT devices. SOD was built to provide a common infrastructure for computer vision applications and to accelerate the use of machine perception in open source as well as commercial products. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • Build Agents and Models on One Platform Icon
    Build Agents and Models on One Platform

    Everything you need to build production-ready agents and models. Access 200+ Google and third-party AI models and tools.

    Gemini Enterprise Agent Platform is Google Cloud's comprehensive platform for developers to build, scale, govern, and optimize agents and models. Choose from Google's most advanced models and third-party models like Anthropic's Claude Model Family.
    Start Free
  • 5
    FastAI.jl

    FastAI.jl

    Repository of best practices for deep learning in Julia

    ...It equips you with reusable components for every part of your project while remaining customizable at every layer. FastAI.jl comes with support for common computer vision and tabular data learning tasks, with more to come.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 6
    TuiCss

    TuiCss

    Text-based user interface CSS library

    ...This kind of interface is very legible because the ultra-contrast colors used and because the reduced effects used on the components in the view. The base of this project is Turbo Vision Framework, but some other frameworks were also checked to introduce some features to TuiCss, like curses, ncurses, Newt, etc. Check the examples page in the wiki to stay on top of some creations, or check the getting started page to start using this library! To start to use TuiCss in your project, you can just download the repository content and import the files that are in the dist folder with the html directives. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 7
    HARFANG 3D engine

    HARFANG 3D engine

    HARFANG 3D source code public repository

    ...Its software suite is tailored to developers, designers and engineers aiming to efficiently and seamlessly develop, implement & deploy 3D solutions (HMI, VR/AR, simulation, interactive 3D), regardless of development language or platform constraints. HARFANG Studio is the ideal 3D editor for creating real-time scenes & animations that match your design vision. It can manage the entire 3D graphics production workflow in a simple and optimized manner, without compromising the integration in other development environments. HARFANG Studio’s philosophy is in line with that of HARFANG 3D engine, a compliant, straightforward, fast & lightweight. Everything that runs in HARFANG Studio is compatible with our Framework and its supported coding languages. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 8
    touch-os

    touch-os

    A friendly OS for touchscreens

    touch OS LTS is a Debian-Gnome derivative, designed with a personal touch to provide a more user-friendly experience, while being optimized for touchscreens. It runs on the stable LTS 6.1 kernel (download from 'Files'). touch OS LIQUORIX is powered by the advanced Liquorix 6.5 kernel, offering improved performance and compatibility with modern computers, especially for multitasking and gaming. Both the LTS and LIQUORIX versions allow you to switch between Wayland and Xorg window...
    Downloads: 4 This Week
    Last Update:
    See Project
  • 9
    OpenEagle

    OpenEagle

    Synthetic Vision/Sectional Maps for experimental aircraft.

    OpenEagle provides interfacing to Flightgear to create a synthetic view of the flight using a Garmin G5 EFIS, and other instruments to display the flight at 10hz. The system also includes tools to make and use browser based moving maps from the FAA sectional "tif" files free to download. OpenEagle replaces NAV, and FIX databases in Flightgear with current FAA data. FAA Special rules airspace such as Class B,C, and D, and MOA's et al are displayed in the synthetic view using translucent...
    Downloads: 0 This Week
    Last Update:
    See Project
  • Save Up to 91% on Cloud Compute With Spot VMs Icon
    Save Up to 91% on Cloud Compute With Spot VMs

    Automatic sustained-use discounts. One free VM per month. No negotiation needed.

    Run batch jobs at 60-91% off with Spot VMs. Long-running workloads get automatic discounts with sustained use.
    Start Free
  • 10
    DataGym.ai

    DataGym.ai

    Open source annotation and labeling tool for image and video assets

    ...AI-assisted annotation tools reduce manual labeling effort, give you more time to finetune ML models and speed up your go to market of new products. Accelerate your computer vision projects by cutting down data preparation time up to 50%. A machine learning model is only as good as its training data. DATAGYM is an end-to-end workbench to create, annotate, manage, and export the right training data for your computer vision models. Your image data can be imported into DATAGYM from your local machine, from any public image URL or directly from an AWS cloud S3 bucket. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 11
    iJEPA

    iJEPA

    Official codebase for I-JEPA

    ...This objective sidesteps generative pixel losses and avoids heavy negative sampling, producing features that transfer strongly with linear probes and minimal fine-tuning. The design scales naturally with Vision Transformer backbones and flexible masking strategies, and it trains stably at large batch sizes. i-JEPA’s predictions are made in embedding space, which is computationally efficient and better aligned with downstream discrimination tasks. The repository provides training recipes, data pipelines, and evaluation code that clarify which masking patterns and architectural choices matter most.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 12
    DeepMind Research

    DeepMind Research

    Implementations and code to accompany DeepMind publications

    This repository collects reference implementations and illustrative code accompanying a wide range of DeepMind publications, making it easier for the research community to reproduce results, inspect algorithms, and build on prior work. The top level organizes many paper-specific directories across domains such as deep reinforcement learning, self-supervised vision, generative modeling, scientific ML, and program synthesis—for example BYOL, Perceiver/Perceiver IO, Enformer for genomics, MeshGraphNets for physics, RL Unplugged, Nowcasting for weather, and more. Each project folder typically includes its own README, scripts, and notebooks so you can run experiments or explore models in isolation, and many link to associated datasets or external environments like DeepMind Lab and StarCraft II. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 13
    Fiducial Objects

    Fiducial Objects

    Multiplanar customizable fiducial markers

    ...Overview: Fiducial Objects is an innovative and versatile project designed to simplify and enhance fiducial marker usage in various applications. Whether you're working in robotics, augmented reality, computer vision, or any other field that requires precise object tracking, our customizable multiplanar fiducial markers are here to meet your needs. Web: https://www.uco.es/investiga/grupos/ava/portfolio/fiducial-object/
    Downloads: 0 This Week
    Last Update:
    See Project
  • 14

    IIDC Camera Control Library

    Capture and control API for IIDC compliant cameras

    ...Besides capture and control, libdc1394 provides a full set of colour space conversion functions (including RAW decoding), vendor specific functions and direct camera register access. Keywords: ieee1394, IIDC, DCAM, firewire, USB, machine vision, computer vision, video capture, library
    Leader badge
    Downloads: 161 This Week
    Last Update:
    See Project
  • 15
    PIFuHD

    PIFuHD

    High-Resolution 3D Human Digitization from A Single Image

    PIFuHD (Pixel-Aligned Implicit Function for 3D human reconstruction at high resolution) is a method and codebase to reconstruct high-fidelity 3D human meshes from a single image. It extends prior PIFu work by increasing resolution and detail, enabling fine geometry in cloth folds, hair, and subtle surface features. The method operates by learning an implicit occupancy / surface function conditioned on the image and camera projection; at inference time it queries dense points to reconstruct a...
    Downloads: 7 This Week
    Last Update:
    See Project
  • 16
    Chameleon LLM

    Chameleon LLM

    Codes for "Chameleon: Plug-and-Play Compositional Reasoning

    Discover Chameleon, our cutting-edge compositional reasoning framework designed to enhance large language models (LLMs) and overcome their inherent limitations, such as outdated information and lack of precise reasoning. By integrating various tools such as vision models, web search engines, Python functions, and rule-based modules, Chameleon delivers more accurate, up-to-date, and precise responses, making it a game-changer in the natural language processing landscape. With GPT-4 at its core, Chameleon has showcased exceptional improvements in accuracy on benchmark tasks, outperforming competitors and setting new industry standards.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 17
    ClassyVision

    ClassyVision

    An end-to-end PyTorch framework for image and video classification

    Classy Vision is a PyTorch-based framework designed for large-scale training and deployment of state-of-the-art image and video classification models. Developed by Facebook Research, it serves as an end-to-end system that simplifies the process of training at scale, reducing redundancy and friction in moving from research to production. Unlike traditional computer vision libraries that focus solely on modular components, Classy Vision provides a complete and unified framework, featuring distributed training, reproducible experiments, and flexible configuration tools. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 18
    T81 558

    T81 558

    Applications of Deep Neural Networks

    ...This course will introduce the student to classic neural network structures, Convolution Neural Networks (CNN), Long Short-Term Memory (LSTM), Gated Recurrent Neural Networks (GRU), General Adversarial Networks (GAN) and reinforcement learning. Application of these architectures to computer vision, time series, security, natural language processing (NLP), and data generation will be covered. High-Performance Computing (HPC) aspects will demonstrate how deep learning can be leveraged both on graphical processing units (GPUs), as well as grids.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 19
    ToMe (Token Merging)

    ToMe (Token Merging)

    A method to increase the speed and lower the memory footprint

    ToMe (Token Merging) is a PyTorch-based optimization framework designed to significantly accelerate Vision Transformer (ViT) architectures without retraining. Developed by researchers at Facebook (Meta AI), ToMe introduces an efficient technique that merges similar tokens within transformer layers, reducing redundant computation while preserving model accuracy. This approach differs from token pruning, which removes background tokens entirely; instead, ToMe merges tokens based on feature similarity, allowing it to compress both foreground and background information efficiently. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 20
    FFCV

    FFCV

    Fast Forward Computer Vision (and other ML workloads!)

    ffcv is a drop-in data loading system that dramatically increases data throughput in model training. From gridding to benchmarking to fast research iteration, there are many reasons to want faster model training. Below we present premade codebases for training on ImageNet and CIFAR, including both (a) extensible codebases and (b) numerous premade training configurations.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 21
    Self-learning-Computer-Science

    Self-learning-Computer-Science

    Resources to learn computer science in your spare time

    Self-learning Computer Science is a curated, open-source guide repository designed to help learners independently study computer science topics using high-quality university-level resources. The author (an undergraduate CS student) assembled links to courses from institutions like MIT, UC Berkeley, Stanford, etc., covering mathematics, programming, data structures/algorithms, computer architecture, machine learning, software engineering and more. It’s aimed at learners who find traditional...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 22
    Robotics Lab

    Robotics Lab

    Open Source from the Robotics Lab research group @ UC3M

    Welcome to the open branch of the Robotics Lab research group @ Universidad Carlos III de Madrid (UC3M). We currently host 3 main projects here: * ASIBOT open source software, which includes basic simulation, control and vision: http://roboticslab.sourceforge.net/asibot * Datasets we use for machine learning: https://sourceforge.net/projects/roboticslab/files/Datasets * Nicolas Burrus' RGBDemo stuff.
    Leader badge
    Downloads: 73 This Week
    Last Update:
    See Project
  • 23
    ImageAI

    ImageAI

    A python library built to empower developers

    ImageAI is an easy-to-use Computer Vision Python library that empowers developers to easily integrate state-of-the-art Artificial Intelligence features into their new and existing applications and systems. It is used by thousands of developers, students, researchers, tutors and experts in corporate organizations around the world. You will find features supported, links to official documentation as well as articles on ImageAI.
    Downloads: 8 This Week
    Last Update:
    See Project
  • 24
    AlgoNotes

    AlgoNotes

    Summary of articles from the official account of Qianmeng Study Notes

    ...Major sections cover ranking and conversion prediction, recall and matching, user profiles, and feature engineering. It also includes recommendation and search topics, computational advertising, big data, and graph algorithms. Additional material covers NLP, computer vision, technical discussions, and job interviews. The repository serves as a structured reference hub for practitioners and learners exploring applied machine learning systems.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 25
    giv
    giv, The G(reat|tk|NU) Image Viewer, is a 8-bit to 32-bit and floating point image and vector viewer for gtk+. It is designed especially for scientific vision and computational geometry and has support for drawing vector graphics on top of an image.
    Leader badge
    Downloads: 2 This Week
    Last Update:
    See Project