Showing 800 open source projects for "vision"

View related business solutions
  • Ship Agents Faster Icon
    Ship Agents Faster

    Transform your applications and workflows into powerful agentic systems at global scale.

    Gemini Enterprise Agent Platform lets you rapidly build, scale, govern and optimize production-ready agents grounded in your organization's data. The platform enables developers to build custom or pre-built agents for virtually any use case. New customers get $300 in free credits.
    Start Free
  • Go from Code to Production URL in Seconds Icon
    Go from Code to Production URL in Seconds

    Cloud Run deploys apps in any language instantly. Scales to zero. Pay only when code runs.

    Skip the Kubernetes configs. Cloud Run handles HTTPS, scaling, and infrastructure automatically. Two million requests free per month.
    Start Free
  • 1
    The Data Fusion Peer is a multitier computer vision internet application. The system provides image processing, motion tracking, and visualization information. Application will convert data into 3-Deminsional and other digital environments.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 2

    Savant

    Python Computer Vision & Video Analytics Framework With Batteries Incl

    Savant is an open-source, high-level framework for building real-time, streaming, highly efficient multimedia AI applications on the Nvidia stack. It helps to develop dynamic, fault-tolerant inference pipelines that utilize the best Nvidia approaches for data center and edge accelerators. Savant is built on DeepStream and provides a high-level abstraction layer for building inference pipelines. It is designed to be easy to use, flexible, and scalable. It is a great choice for building...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 3
    eye-pointer is a set of libraries and programs for computer vision, especially for locating the spot on the screen you're looking at. The ultimate goal is to incorporate that facility into a pointer (mouse) driver.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 4
    The project is aimed at automatic target following using a camera , a computer vision system and a microcontroller that moves the cam. The project should mainly work under linux and it might be ported into windows,
    Downloads: 0 This Week
    Last Update:
    See Project
  • Save Up to 91% on Cloud Compute With Spot VMs Icon
    Save Up to 91% on Cloud Compute With Spot VMs

    Automatic sustained-use discounts. One free VM per month. No negotiation needed.

    Run batch jobs at 60-91% off with Spot VMs. Long-running workloads get automatic discounts with sustained use.
    Start Free
  • 5
    ViKi (Virtual Interactive keyboard Interface) is a global framework that enables contactless human machine interaction using computer vision techniques. Only a simple webcam is sufficient to emulate traditional devices such as mouse and keyboard do.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 6
    Ministral 3 8B Base 2512

    Ministral 3 8B Base 2512

    Versatile 8B-base multimodal LLM, flexible foundation for custom AI

    ...Because it comes from the edge-optimized Ministral 3 family, it remains deployable on reasonably powerful hardware while offering a good balance between capability and resource use. Its multilingual and multimodal pretraining enables broad applicability across languages and tasks — from generation to classification to vision-language tasks.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 7
    uurt-humanoid

    uurt-humanoid

    Source Code for UURT Humanoid KidSize robots

    This project contains main codes and documentations of controller program of Urmia University Robotic Team (UURT) KidSize/Humanoid robots.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 8
    Qwen3.6-35B-A3B

    Qwen3.6-35B-A3B

    Open multimodal model for coding, agents, and long-context tasks

    ...Architecturally, it uses a Mixture-of-Experts design with 35B total parameters and 3B active, supports a native 262K-token context window, and can be extended to about 1M tokens with YaRN. It also performs strongly across coding, agent, vision, reasoning, and document-understanding benchmarks.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 9
    Devstral Small 2

    Devstral Small 2

    Lightweight 24B agentic coding model with vision and long context

    ...With 24B parameters and FP8 instruct tuning, it delivers strong instruction following while remaining lightweight enough for local and on-device deployment. The model achieves competitive performance on SWE-bench, validating its effectiveness for real-world coding and automation tasks. It introduces vision capabilities, enabling image understanding alongside text for more versatile development workflows. Devstral Small 2 supports a 256k context window, allowing it to reason across large repositories, long diffs, and extended technical contexts. Its architecture improves generalization across diverse prompts and coding environments while leveraging advanced attention scaling techniques.
    Downloads: 0 This Week
    Last Update:
    See Project
  • $300 Free Credits to Build on Google Cloud Icon
    $300 Free Credits to Build on Google Cloud

    New customers can spin up VMs, build with AI, and query data at no cost.

    Put your $300 in credit toward real workloads, then keep building with free monthly usage for 20+ products. No commitment and no charge until you upgrade.
    Start Free
  • 10
    OpenVLA 7B

    OpenVLA 7B

    Vision-language-action model for robot control via images and text

    OpenVLA 7B is a multimodal vision-language-action model trained on 970,000 robot manipulation episodes from the Open X-Embodiment dataset. It takes camera images and natural language instructions as input and outputs normalized 7-DoF robot actions, enabling control of multiple robot types across various domains. Built on top of LLaMA-2 and DINOv2/SigLIP visual backbones, it allows both zero-shot inference for known robot setups and parameter-efficient fine-tuning for new domains. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 11
    CLIP-ViT-bigG-14-laion2B-39B-b160k

    CLIP-ViT-bigG-14-laion2B-39B-b160k

    CLIP ViT-bigG/14: Zero-shot image-text model trained on LAION-2B

    CLIP-ViT-bigG-14-laion2B-39B-b160k is a powerful vision-language model trained on the English subset of the LAION-5B dataset using the OpenCLIP framework. Developed by LAION and trained by Mitchell Wortsman on Stability AI’s compute infrastructure, it pairs a ViT-bigG/14 vision transformer with a text encoder to perform contrastive learning on image-text pairs. This model excels at zero-shot image classification, image-to-text and text-to-image retrieval, and can be adapted for tasks such as image captioning or generation guidance. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 12
    MiMo-V2.6-Flash

    MiMo-V2.6-Flash

    Efficient 309B omnimodal MoE for coding, agents, vision, and audio

    ...Training uses a unified mixed RL process rather than separate domain-specific runs, alongside asynchronous GRPO and groupwise agentic grading that rewards higher-quality and more efficient solutions. Its architecture combines sliding-window and global attention with a 681M-parameter vision encoder and dedicated audio encoders. A five-layer speculative decoder predicts multiple subsequent tokens to accelerate inference.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 13
    MiMo-V2.6-Pro

    MiMo-V2.6-Pro

    1T omnimodal MoE model for coding, agents, and long-horizon reasoning

    MiMo-V2.6-Pro is Xiaomi MiMo’s flagship open-weight omnimodal model, built to scale reinforcement learning toward self-improvement across coding, agents, vision, and cybersecurity. Its sparse Mixture-of-Experts architecture contains 1.02T total parameters with 42B activated per token, using 384 routed experts with eight active per token. The model natively processes text, images, video, and audio and supports a 1M-token context window for large repositories, extended tool traces, and multi-session agent workflows. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 14
    Ministral 3 3B Base 2512

    Ministral 3 3B Base 2512

    Small 3B-base multimodal model ideal for custom AI on edge hardware

    Ministral 3 3B Base 2512 is the smallest model in the Ministral 3 family, offering a compact yet capable multimodal architecture suited for lightweight AI applications. It combines a 3.4B-parameter language model with a 0.4B vision encoder, enabling both text and image understanding in a tiny footprint. As the base pretrained model, it is not fine-tuned for instructions or reasoning, making it the ideal foundation for custom post-training, domain adaptation, or specialized downstream tasks. The model is fully optimized for edge deployment and can run locally on a single GPU, fitting in 16GB VRAM in BF16 or less than 8GB when quantized. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 15
    Ministral 3 3B Reasoning 2512

    Ministral 3 3B Reasoning 2512

    Compact 3B-param multimodal model for efficient on-device reasoning

    Ministral 3 3B Reasoning 2512 is the smallest reasoning-capable model in the Ministal-3 family, yet delivers a surprisingly capable multimodal and multilingual base for lightweight AI applications. It pairs a 3.4B-parameter language model with a 0.4B-parameter vision encoder, enabling it to understand both text and image inputs. This reasoning-tuned variant is optimized for tasks like math, coding, and other STEM-related problem solving, making it suitable for applications that require logical reasoning, analysis, or structured thinking. Despite its modest size, the model is designed for edge deployment and can run locally, fitting in ~16 GB of VRAM in BF16 or under 8 GB of RAM/VRAM when quantized. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 16
    Ministral 3 14B Base 2512

    Ministral 3 14B Base 2512

    Powerful 14B-base multimodal model — flexible base for fine-tuning

    Ministral 3 14B Base 2512 is the largest model in the Ministral 3 line, offering state-of-the-art language and vision capabilities in a dense, base-pretrained form. It combines a 13.5B-parameter language model with a 0.4B-parameter vision encoder, enabling both high-quality text understanding/generation and image-aware tasks. As a “base” model (i.e. not fine-tuned for instruction or reasoning), it provides a flexible foundation ideal for custom fine-tuning or downstream specialization. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 17
    Control application for robotics. Vision, Trajectory, IA and Remote control.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 18
    Qwen3.6-27B

    Qwen3.6-27B

    Dense multimodal Qwen model for coding, agents, and long context

    Qwen3.6-27B is an open-weight multimodal model built to deliver strong real-world coding, agent, and long-context performance in a dense 27B-parameter architecture. It combines a causal language model with a vision encoder and supports text, image, and video inputs, making it suitable for both software workflows and broader multimodal tasks. The model emphasizes stability and practical developer utility, with major improvements in agentic coding, frontend generation, and repository-level reasoning. It also introduces thinking preservation, allowing it to retain reasoning traces from earlier turns to improve consistency, reduce repeated computation, and support iterative agent workflows. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 19
    Qwen2.5-VL-3B-Instruct

    Qwen2.5-VL-3B-Instruct

    Qwen2.5-VL-3B-Instruct: Multimodal model for chat, vision & video

    Qwen2.5-VL-3B-Instruct is a 3.75 billion parameter multimodal model by Qwen, designed to handle complex vision-language tasks in both image and video formats. As part of the Qwen2.5 series, it supports image-text-to-text generation with capabilities like chart reading, object localization, and structured data extraction. The model can serve as an intelligent visual agent capable of interacting with digital interfaces and understanding long-form videos by dynamically sampling resolution and frame rate. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 20
    Alpamayo 2 Super

    Alpamayo 2 Super

    Open VLA model for autonomous driving reasoning and planning

    Alpamayo2-Super is NVIDIA’s flagship open reasoning Vision-Language-Action (VLA) model for autonomous vehicle development, designed to accelerate Level 4 robotaxi and autonomous driving systems. Built on the NVIDIA Cosmos platform, it enables vehicles to perceive, reason, plan, and generate driving actions using human-like decision making rather than relying solely on predefined rules. The model processes multi-camera video, navigation signals, and driving context to produce driving trajectories alongside interpretable Chain-of-Causation reasoning, improving transparency for validation and safety analysis. ...
    Downloads: 1 This Week
    Last Update:
    See Project
  • 21
    Inkling-Small

    Inkling-Small

    Efficient multimodal MoE model for coding, tools, and reasoning

    ...Its 42-layer decoder routes each token through six of 256 specialized experts plus two shared experts, while hybrid local and global attention supports efficient processing. Inkling-Small performs strongly across software engineering, tool use, mathematics, vision, and audio benchmarks, including 80.2% on SWE-Bench Verified and 95.5% on AIME 2026.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 22
    MiMo-V2.5

    MiMo-V2.5

    Omnimodal AI model for agents, coding, and long-context tasks

    ...MiMo-V2.5 delivers near-Pro-level performance in coding, reasoning, and agent tasks while maintaining lower cost and faster inference speeds. It also integrates advanced components such as multi-token prediction modules and specialized vision and audio encoders, making it well-suited for autonomous agents and software development.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 23
    Ministral 3 8B Instruct 2512

    Ministral 3 8B Instruct 2512

    Compact 8B multimodal instruct model optimized for edge deployment

    Ministral 3 8B Instruct 2512 is a balanced, efficient model in the Ministral 3 family, offering strong multimodal capabilities within a compact footprint. It combines an 8.4B-parameter language model with a 0.4B vision encoder, enabling both text reasoning and image understanding. This FP8 instruct-fine-tuned variant is optimized for chat, instruction following, and structured outputs, making it ideal for daily assistant tasks and lightweight agentic workflows. Designed for edge deployment, the model can run on a wide range of hardware and fits locally on a single 12GB GPU, with the option for even smaller quantized configurations. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 24
    A collection of tools and code for a stereoscopic responsive workbench. Settled in Boğaziçi University, Istanbul, our Media Lab is a small group of 3D Computer Graphics and Vision enthusiast students from all degrees.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 25
    Mistral Large 3 675B Base 2512

    Mistral Large 3 675B Base 2512

    Frontier-scale 675B multimodal base model for custom AI training

    ...The model is engineered for reliability, long-context comprehension, and stable performance across many enterprise, scientific, and knowledge-intensive workloads. Its architecture includes a powerful language MoE and a 2.5B-parameter vision encoder, enabling multimodal understanding out of the box. Mistral Large 3 Base supports deployment on-premises using FP8 or NVFP4 formats, enabling high-performance workflows on B200, H200, H100, or A100 hardware.
    Downloads: 0 This Week
    Last Update:
    See Project