Showing 440 open source projects for "benchmarks"

View related business solutions
  • Ship Agents Faster Icon
    Ship Agents Faster

    Transform your applications and workflows into powerful agentic systems at global scale.

    Gemini Enterprise Agent Platform lets you rapidly build, scale, govern and optimize production-ready agents grounded in your organization's data. The platform enables developers to build custom or pre-built agents for virtually any use case. New customers get $300 in free credits.
    Start Free
  • Veeam Data Platform v13.1 - Get Your Free Trial Icon
    Veeam Data Platform v13.1 - Get Your Free Trial

    Secure by design, portable by default. Recover clean, fast, anywhere. Start a free trial.

    Try Veeam Data Platform today. Experience the unified platform that's secure by design, portable by default, and proven to recover clean, fast, and anywhere.
    Try it Free
  • 1
    Qwen3.8-2.4T-A95B

    Qwen3.8-2.4T-A95B

    Massive 2.4T MoE model for coding, agents, research, and reasoning

    ...It is a text-only, thinking-first model and supports deployment through vLLM, SGLang, and TokenSpeed, with strong results across coding, tool use, research, and professional benchmarks.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 2
    Inkling-Small

    Inkling-Small

    Efficient multimodal MoE model for coding, tools, and reasoning

    ...Its 42-layer decoder routes each token through six of 256 specialized experts plus two shared experts, while hybrid local and global attention supports efficient processing. Inkling-Small performs strongly across software engineering, tool use, mathematics, vision, and audio benchmarks, including 80.2% on SWE-Bench Verified and 95.5% on AIME 2026.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 3
    Inkling

    Inkling

    Frontier multimodal MoE model for coding and AI agent workflows

    ...Trained from scratch on approximately 45 trillion multimodal tokens, Inkling introduces controllable reasoning effort, allowing users to trade off latency and reasoning depth depending on the task. It is optimized for software engineering, tool use, and large-scale autonomous workflows, with strong performance on coding and agent benchmarks.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 4
    Nex-N2-mini

    Nex-N2-mini

    Compact agentic model for coding, tools, and productivity tasks

    ...Nex-N2-mini supports image-text-to-text workflows, explicit reasoning traces, robust function calling, and deployment through Transformers, vLLM, SGLang, Docker, and quantized local apps. It performs strongly across agentic, coding, search, and reasoning benchmarks, including SWE-Bench, Terminal-Bench, BrowseComp, Toolathlon, and GPQA.
    Downloads: 0 This Week
    Last Update:
    See Project
  • $300 Free Credits to Build on Google Cloud Icon
    $300 Free Credits to Build on Google Cloud

    New customers can spin up VMs, build with AI, and query data at no cost.

    Put your $300 in credit toward real workloads, then keep building with free monthly usage for 20+ products. No commitment and no charge until you upgrade.
    Start Free
  • 5
    Nex-N2-Pro

    Nex-N2-Pro

    Large agentic model for coding, tools, research, and execution

    ...It supports image-text-to-text workflows, explicit reasoning traces, robust function calling, and deployment through Transformers, vLLM, SGLang, Docker, and quantized local apps. Nex-N2-Pro performs strongly across agentic, coding, search, and reasoning benchmarks, including Terminal-Bench, SWE-Bench Pro, BrowseComp, Toolathlon, WideSearch, GPQA Diamond, and GDPval.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 6
    Gemma 4

    Gemma 4

    Google’s flagship dense multimodal model for coding and reasoning

    ...Google positions it as a frontier-level model that can run on consumer GPUs and workstations while achieving leading results across reasoning, mathematics, coding, and multimodal benchmarks.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 7
    GigaChat 3 Ultra

    GigaChat 3 Ultra

    High-performance MoE model with MLA, MTP, and multilingual reasoning

    ...Its training corpus incorporates ten languages, enriched with books, academic sources, code datasets, mathematical tasks, and more than 5.5 trillion tokens of high-quality synthetic data. This combination significantly boosts reasoning, coding, and multilingual performance across modern benchmarks. Designed for high-performance deployment, GigaChat 3 Ultra supports major inference engines and offers optimized BF16 and FP8 execution paths for cluster-grade hardware.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 8
    DeepSeek-V3.2-Speciale

    DeepSeek-V3.2-Speciale

    High-compute ultra-reasoning model surpassing model surpassing GPT-5

    DeepSeek-V3.2-Speciale is the high-compute, ultra-reasoning variant of DeepSeek-V3.2, designed specifically to push the boundaries of mathematical, logical, and algorithmic intelligence. It builds on the DeepSeek Sparse Attention (DSA) framework, delivering dramatically improved long-context efficiency while preserving full model quality. Unlike the standard version, Speciale is tuned exclusively for deep reasoning and therefore does not support tool-calling, focusing its full capacity on...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 9
    Qwen3-Next

    Qwen3-Next

    Qwen3-Next: 80B instruct LLM with ultra-long context up to 1M tokens

    ...Multi-Token Prediction (MTP) boosts both training and inference, while stability optimizations such as weight-decayed and zero-centered layernorm ensure robustness. Benchmarks show it performs comparably to larger models like Qwen3-235B on reasoning, coding, multilingual, and alignment tasks while requiring only a fraction of the training cost.
    Downloads: 0 This Week
    Last Update:
    See Project
  • Go from Code to Production URL in Seconds Icon
    Go from Code to Production URL in Seconds

    Cloud Run deploys apps in any language instantly. Scales to zero. Pay only when code runs.

    Skip the Kubernetes configs. Cloud Run handles HTTPS, scaling, and infrastructure automatically. Two million requests free per month.
    Start Free
  • 10
    Mellum-4b-base

    Mellum-4b-base

    JetBrains’ 4B parameter code model for completions

    ...While the base model is not fine-tuned for downstream tasks, it is designed to be easily adapted through supervised fine-tuning (SFT) or reinforcement learning (RL). Benchmarks on RepoBench, SAFIM, and HumanEval demonstrate its competitive performance, with specialized fine-tuned versions for Python already showing strong improvements.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 11
    Llama-3.2-1B-Instruct

    Llama-3.2-1B-Instruct

    Instruction-tuned 1.2B LLM for multilingual text generation by Meta

    ...Llama-3.2-1B is lightweight enough for deployment on constrained devices like smartphones, using formats like SpinQuant and QLoRA to reduce model size and latency. Despite its small size, it performs competitively across benchmarks such as MMLU, ARC, and TLDR summarization. The model is distributed under the Llama 3.2 Community License, requiring attribution and adherence to Meta’s Acceptable Use Policy.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 12
    fashion-clip

    fashion-clip

    CLIP model fine-tuned for zero-shot fashion product classification

    ...FashionCLIP 2.0, the latest version, uses the laion/CLIP-ViT-B-32-laion2B-s34B-b79K checkpoint for improved accuracy, achieving better F1 scores across multiple benchmarks compared to earlier versions. It supports multilingual fashion queries and works best with clean, product-style images against white backgrounds. The model can be used for product search, recommendation systems, or visual tagging in e-commerce platforms.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 13
    Laguna M.1

    Laguna M.1

    Flagship Poolside model for agentic coding and software engineering

    ...Trained from scratch on roughly 30 trillion tokens using Poolside’s in-house “Model Factory” pipeline, the model focuses on complex software development tasks, repository-scale reasoning, tool use, and multi-step agent execution. Laguna M.1 was designed to compete with leading frontier coding models on benchmarks such as SWE-Bench, Terminal-Bench, and other agentic engineering evaluations. It supports reasoning, tool calling, and long-context workflows, making it suitable for autonomous coding agents, software maintenance, debugging, and large-scale development projects.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 14
    VaultGemma

    VaultGemma

    VaultGemma: 1B DP-trained Gemma variant for private NLP tasks

    ...Training ran on TPU v6e using JAX and Pathways with privacy-preserving algorithms (DP-SGD, truncated Poisson subsampling) and DP scaling laws to balance compute and privacy budgets. Benchmarks on the 1B pre-trained checkpoint show expected utility trade-offs (e.g., HellaSwag 10-shot 39.09, BoolQ 0-shot 62.04, PIQA 0-shot 68.00), reflecting its privacy-first design.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 15
    Turn-key secure credit card processing appliance
    Downloads: 0 This Week
    Last Update:
    See Project