Showing 17 open source projects for "flash image software"

View related business solutions
  • Demo Series - Small Business Backup By Veeam Icon
    Demo Series - Small Business Backup By Veeam

    Learn how to protect your Microsoft 365 data, with simple, actionable tips today.

    Watch this on-demand demo series and learn how to protect your Microsoft 365 data with clear, simple, actionable steps that are easy to implement for businesses of all sizes.
    Watch Demo Series
  • MongoDB Atlas runs apps anywhere Icon
    MongoDB Atlas runs apps anywhere

    Deploy in 115+ regions with the modern database for every enterprise.

    MongoDB Atlas gives you the freedom to build and run modern applications anywhere—across AWS, Azure, and Google Cloud. With global availability in over 115 regions, Atlas lets you deploy close to your users, meet compliance needs, and scale with confidence across any geography.
    Start Free
  • 1
    DSH Web UI

    DSH Web UI

    Plugin and skin collection for DeepSeek Harness (DSH) Web UI

    DSH Web UI is a modular plugin and skin collection that expands the DeepSeek Harness browser interface into a broader development workspace. Its components attach through the official DSH profile mechanism without modifying the Harness source code. Users can add a task board that executes real agent sessions and supports scheduled jobs through cron expressions. A Git graph visualizes branches and commit history, while additional panels provide files, editing, terminals, Git operations, and...
    Downloads: 5 This Week
    Last Update:
    See Project
  • 2
    Stable Diffusion Web UI Extensions

    Stable Diffusion Web UI Extensions

    Extension index for stable-diffusion-webui

    This repository serves as the official index used by the Stable Diffusion Web UI to discover and install extensions. It aggregates metadata for hundreds of community plugins—image utilities, ControlNet tools, upscalers, prompt helpers, animation suites—so users can browse and add capabilities directly from the UI. The index maintains short descriptions, tags, and repository links, enabling quick filtering by purpose or workflow. It also standardizes submission format so extension authors can...
    Downloads: 4 This Week
    Last Update:
    See Project
  • 3
    GLM-4.6V

    GLM-4.6V

    GLM-4.6V/4.5V/4.1V-Thinking, towards versatile multimodal reasoning

    GLM-4.6V represents the latest generation of the GLM-V family and marks a major step forward in multimodal AI by combining advanced vision-language understanding with native “tool-call” capabilities, long-context reasoning, and strong generalization across domains. Unlike many vision-language models that treat images and text separately or require intermediate conversions, GLM-4.6V allows inputs such as images, screenshots or document pages directly as part of its reasoning pipeline — and...
    Downloads: 1 This Week
    Last Update:
    See Project
  • 4
    CO3D (Common Objects in 3D)

    CO3D (Common Objects in 3D)

    Tooling for the Common Objects In 3D dataset

    CO3Dv2 (Common Objects in 3D, version 2) is a large-scale 3D computer vision dataset and toolkit from Facebook Research designed for training and evaluating category-level 3D reconstruction methods using real-world data. It builds upon the original CO3Dv1 dataset, expanding both scale and quality—featuring 2× more sequences and 4× more frames, with improved image fidelity, more accurate segmentation masks, and enhanced annotations for object-centric 3D reconstruction. CO3Dv2 enables research...
    Downloads: 1 This Week
    Last Update:
    See Project
  • Save Up to 91% on Cloud Compute With Spot VMs Icon
    Save Up to 91% on Cloud Compute With Spot VMs

    Automatic sustained-use discounts. One free VM per month. No negotiation needed.

    Run batch jobs at 60-91% off with Spot VMs. Long-running workloads get automatic discounts with sustained use.
    Start Free
  • 5
    CycleGAN

    CycleGAN

    Software that can generate photos from paintings

    CycleGAN — in its original form — is a landmark in deep learning for image-to-image translation without paired data. Rather than requiring matching image pairs between source and target domains (which are often hard or impossible to obtain), CycleGAN learns two mappings — one from domain A to B, and another back from B to A — along with a cycle-consistency loss that encourages the round-trip to reconstruct the original image. This innovation lets the model learn domain-to-domain translations...
    Downloads: 1 This Week
    Last Update:
    See Project
  • 6
    Smaug Flash

    Smaug Flash

    304B MoE model optimized for agentic coding and long-running workflows

    Smaug Flash is an open-weight agentic coding model from Abacus.AI, fine-tuned from DeepSeek-V4-Flash-0731 to improve autonomous software engineering, tool use, and long-running agent workflows. It uses a 304B-parameter Mixture-of-Experts architecture with 43 layers, 256 routed experts, six selected experts per token, and one shared expert. Its attention system combines Multi-Head Latent Attention with a sparse token indexer, while DSpark multi-token prediction provides speculative decoding for faster generation. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 7
    GLM-5.3-Flash

    GLM-5.3-Flash

    Efficient 320B multimodal MoE model for coding and autonomous agents

    ...It also uses Manifold-Constrained Hyper-Connections (mHC) to improve scaling efficiency and was pretrained on a 30-trillion-token multimodal corpus. GLM-5.3-Flash supports text and image inputs and is particularly optimized for coding and autonomous agent workloads, approaching larger frontier models on related benchmarks while improving over GLM-5.2. It supports local deployment through SGLang, vLLM, TokenSpeed, and KTransformers and is released under the MIT license.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 8
    Qwen3.8-Flash-Next

    Qwen3.8-Flash-Next

    Efficient multimodal MoE model for coding, reasoning, and AI agents

    Qwen3.8-Flash-Next is Qwen’s experimental open-weight multimodal model previewing the architecture planned for Qwen4. It uses 125B language-model parameters with only 6B activated per token, supplemented by 51B n-gram embedding parameters and 4B for multi-token prediction. Its hybrid architecture combines Gated DeltaNet with Qwen Sparse Attention (QSA), which processes micro-blocks rather than individual tokens to reduce latency in long-context agent workloads. The model also introduces...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 9
    MiMo-V2.6-Flash

    MiMo-V2.6-Flash

    Efficient 309B omnimodal MoE for coding, agents, vision, and audio

    MiMo-V2.6-Flash is Xiaomi MiMo’s efficiency-balanced open-weight omnimodal model, designed to scale reinforcement learning across coding, general agents, visual tasks, and cybersecurity. Its sparse Mixture-of-Experts architecture contains 309B total parameters while activating only 15B per token, using 256 routed experts with eight active per token. The model natively processes text, images, video, and audio and supports a 1M-token context window for large repositories, extended tool traces,...
    Downloads: 0 This Week
    Last Update:
    See Project
  • Ship Agents Faster Icon
    Ship Agents Faster

    Transform your applications and workflows into powerful agentic systems at global scale.

    Gemini Enterprise Agent Platform lets you rapidly build, scale, govern and optimize production-ready agents grounded in your organization's data. The platform enables developers to build custom or pre-built agents for virtually any use case. New customers get $300 in free credits.
    Start Free
  • 10
    OpenVLA 7B

    OpenVLA 7B

    Vision-language-action model for robot control via images and text

    OpenVLA 7B is a multimodal vision-language-action model trained on 970,000 robot manipulation episodes from the Open X-Embodiment dataset. It takes camera images and natural language instructions as input and outputs normalized 7-DoF robot actions, enabling control of multiple robot types across various domains. Built on top of LLaMA-2 and DINOv2/SigLIP visual backbones, it allows both zero-shot inference for known robot setups and parameter-efficient fine-tuning for new domains. The model...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 11
    Hy3

    Hy3

    Open code agent for Lean 4 proofs and formal software verification

    ...Leanstral accepts text and image inputs and produces text output, enabling multimodal workflows around mathematics, code, and specifications. It supports configurable reasoning effort, allowing users to disable reasoning or enable high-effort reasoning for complex prompts. This updated version of the original Leanstral focuses on performant, cost-effective formal coding and theorem-proving workflows.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 12
    Leanstral 1.5

    Leanstral 1.5

    Open code agent for Lean 4 proofs and formal software verification

    ...Leanstral accepts text and image inputs and produces text output, enabling multimodal workflows around mathematics, code, and specifications. It supports configurable reasoning effort, allowing users to disable reasoning or enable high-effort reasoning for complex prompts. This updated version of the original Leanstral focuses on performant, cost-effective formal coding and theorem-proving workflows.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 13
    Devstral Small 2

    Devstral Small 2

    Lightweight 24B agentic coding model with vision and long context

    ...It introduces vision capabilities, enabling image understanding alongside text for more versatile development workflows. Devstral Small 2 supports a 256k context window, allowing it to reason across large repositories, long diffs, and extended technical contexts. Its architecture improves generalization across diverse prompts and coding environments while leveraging advanced attention scaling techniques.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 14
    Qwen3.6-27B

    Qwen3.6-27B

    Dense multimodal Qwen model for coding, agents, and long context

    Qwen3.6-27B is an open-weight multimodal model built to deliver strong real-world coding, agent, and long-context performance in a dense 27B-parameter architecture. It combines a causal language model with a vision encoder and supports text, image, and video inputs, making it suitable for both software workflows and broader multimodal tasks. The model emphasizes stability and practical developer utility, with major improvements in agentic coding, frontend generation, and repository-level reasoning. It also introduces thinking preservation, allowing it to retain reasoning traces from earlier turns to improve consistency, reduce repeated computation, and support iterative agent workflows. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 15
    Kimi K2.7 Code

    Kimi K2.7 Code

    Coding-focused Kimi model for long-horizon agent workflows

    ...The model supports image and video input, native INT4 quantization, interleaved thinking, and multi-step tool calling. It also forces preserve-thinking mode by default, retaining full reasoning context across multi-turn interactions to improve coding-agent consistency. K2.7 Code is recommended for use through Kimi Code CLI and can be deployed with vLLM, SGLang, or KTransformers.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 16
    Inkling-Small

    Inkling-Small

    Efficient multimodal MoE model for coding, tools, and reasoning

    Inkling-Small is an open-weight general-purpose multimodal model from Thinking Machines Lab, designed for agentic systems, coding assistants, chatbots, retrieval workflows, and natural-language applications. It accepts text, images, and audio as input and produces text output, with multilingual and multi-programming-language capabilities. The model uses a sparse Mixture-of-Experts architecture with 276B total parameters and 12B active per token, enabling strong performance with lower...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 17
    MiMo-V2.5

    MiMo-V2.5

    Omnimodal AI model for agents, coding, and long-context tasks

    MiMo-V2.5 is a native omnimodal large language model developed by Xiaomi, designed for advanced agentic workflows, multimodal reasoning, and long-context processing. Built on a Mixture-of-Experts architecture with approximately 309B total parameters and around 15B activated per inference, it balances high capability with efficient execution. The model natively processes text, images, video, and audio within a unified system, enabling cross-modal understanding and complex task execution in a...
    Downloads: 0 This Week
    Last Update:
    See Project
  • Previous
  • You're on page 1
  • Next