Showing 440 open source projects for "benchmarks"

View related business solutions
  • Ship Agents Faster Icon
    Ship Agents Faster

    Transform your applications and workflows into powerful agentic systems at global scale.

    Gemini Enterprise Agent Platform lets you rapidly build, scale, govern and optimize production-ready agents grounded in your organization's data. The platform enables developers to build custom or pre-built agents for virtually any use case. New customers get $300 in free credits.
    Start Free
  • $300 Free Credits to Build on Google Cloud Icon
    $300 Free Credits to Build on Google Cloud

    New customers can spin up VMs, build with AI, and query data at no cost.

    Put your $300 in credit toward real workloads, then keep building with free monthly usage for 20+ products. No commitment and no charge until you upgrade.
    Start Free
  • 1
    pydantic

    pydantic

    Data parsing and validation using Python type hints

    Data validation and settings management using Python type hinting. Fast and extensible, pydantic plays nicely with your linters/IDE/brain. Define how data should be in pure, canonical Python 3.6+; validate it with pydantic. id is of type int; the annotation-only declaration tells pydantic that this field is required. Strings, bytes or floats will be coerced to ints if possible; otherwise an exception will be raised. name is inferred as a string from the provided default; because it has a...
    Downloads: 1 This Week
    Last Update:
    See Project
  • 2
    GLM-4.5

    GLM-4.5

    GLM-4.5: Open-source LLM for intelligent agents by Z.ai

    ...Both models unify reasoning, coding, and intelligent agent capabilities, providing two modes: a thinking mode for complex reasoning and tool usage, and a non-thinking mode for immediate responses. They are released under the MIT license, allowing commercial use and secondary development. GLM-4.5 achieves strong performance on 12 industry-standard benchmarks, ranking 3rd overall, while GLM-4.5-Air balances competitive results with greater efficiency. The models support FP8 and BF16 precision, and can handle very large context windows of up to 128K tokens. Flexible inference is supported through frameworks like vLLM and SGLang with tool-call and reasoning parsers included.
    Downloads: 23 This Week
    Last Update:
    See Project
  • 3
    FlashKDA

    FlashKDA

    High-performance Kimi Delta Attention kernels

    FlashKDA is an open-source library of high-performance CUDA kernels for Kimi Delta Attention, implemented on NVIDIA CUTLASS. It is intended to accelerate the forward pass used by KDA-based language models on modern NVIDIA GPUs. The package integrates with flash-linear-attention and can be selected automatically as the backend for chunk_kda during inference. It supports recurrent state input and output, variable-length batches, internal gating, query-key normalization, and beta activation....
    Downloads: 0 This Week
    Last Update:
    See Project
  • 4
    SpaceWasm

    SpaceWasm

    A flight-compliant WebAssembly interpreter

    ...Embedders can expose controlled host functions while restricting program access, compute time, and system interactions. The project emphasizes sandboxing, portability, predictable resource use, and compatibility with flight-software requirements. Tests, fuzzing, benchmarks, requirements, and integration components are included in the Rust codebase.
    Downloads: 0 This Week
    Last Update:
    See Project
  • Fully Managed MySQL, PostgreSQL, and SQL Server Icon
    Fully Managed MySQL, PostgreSQL, and SQL Server

    Automatic backups, patching, replication, and failover. Focus on your app, not your database.

    Cloud SQL handles your database ops end to end, so you can focus on your app.
    Start Free
  • 5
    DeepSpec

    DeepSpec

    A full-stack codebase for training and evaluating speculative decoding

    DeepSpec is a full-stack codebase for training and evaluating draft models used in speculative decoding. It provides the components needed to prepare data, train draft models, and measure acceptance behavior against target models. The workflow starts with data preparation, including prompt download, target answer regeneration, and target cache construction. It then trains a draft model using configuration files for different algorithms and target model setups. The evaluation pipeline...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 6
    Private Mind

    Private Mind

    Your private AI mind. In your pocket

    Private Mind is a mobile AI app built around fully offline, on-device intelligence. It is designed for users who want AI conversations without sending personal data to cloud servers. The app keeps conversations and data on the device, which makes it useful for privacy-sensitive workflows, personal notes, local assistance, and offline use. It supports customizable AI behavior, including choosing supported models, uploading models, and adjusting system prompts. The project also includes...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 7
    whichllm

    whichllm

    Find the local LLM that actually runs and performs best

    ...It detects the machine’s available resources, including GPU, CPU, memory, and storage, then recommends models based on practical fit rather than parameter count alone. The project is useful for users who are unsure which local LLM will perform well on their system. It focuses on real, recency-aware benchmarks so recommendations better reflect current model performance. whichllm is especially helpful for developers, AI hobbyists, and researchers comparing local inference options across NVIDIA, AMD, Apple Silicon, and CPU-only environments. Its main value is reducing guesswork when choosing a local model to download and run.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 8
    Claude Ads

    Claude Ads

    Comprehensive paid advertising audit & optimization skill

    ...It processes user-provided data such as exports or screenshots and evaluates campaigns using hundreds of predefined checks. The system generates structured reports, identifies inefficiencies, and suggests optimization strategies based on industry benchmarks. It supports platforms like Google Ads, Meta Ads, TikTok, LinkedIn, and more, offering a unified analysis workflow. The architecture uses parallel subagents to speed up audits and includes financial modeling and A/B testing guidance. It runs locally, ensuring privacy and control over sensitive advertising data. The project is aimed at replacing manual audit workflows with fast, automated, and repeatable analysis.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 9
    Wasm3

    Wasm3

    A fast WebAssembly interpreter and the most universal WASM runtime

    ...Unlike JIT-based runtimes, Wasm3 uses an interpreter architecture that prioritizes portability, low memory usage, and predictable execution, making it especially suitable for environments where resources are limited. It is widely regarded as one of the fastest interpreters for WebAssembly, achieving competitive performance benchmarks despite not relying on compilation techniques. The project emphasizes universality, meaning it can run consistently across architectures without requiring platform-specific optimizations or dependencies. Wasm3 also integrates cleanly into host applications, allowing developers to embed WebAssembly execution into their systems with minimal overhead.
    Downloads: 0 This Week
    Last Update:
    See Project
  • Go from Code to Production URL in Seconds Icon
    Go from Code to Production URL in Seconds

    Cloud Run deploys apps in any language instantly. Scales to zero. Pay only when code runs.

    Skip the Kubernetes configs. Cloud Run handles HTTPS, scaling, and infrastructure automatically. Two million requests free per month.
    Start Free
  • 10
    NitroGen

    NitroGen

    A Foundation Model for Generalist Gaming Agents

    ...This approach enables the model to control agents in different game genres and contexts, performing tasks that range from complex exploration and combat to fine-grained control in platformers, demonstrating adaptability across unseen environments. The project draws on MineDojo’s broader ecosystem for embodied AI, where multi-modal inputs and richly diverse benchmarks help push toward generalist AI capable of interactive decision making.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 11
    Meta Agents Research Environments (ARE)

    Meta Agents Research Environments (ARE)

    Meta Agents Research Environments is a comprehensive platform

    Meta Agents Research Environments (ARE) is a simulation and benchmarking platform. It is designed to evaluate AI agents in dynamic, evolving, multi-step tasks. Unlike static benchmarks, ARE supports environments where agents must adapt to changes over time and reason over sequences of actions. It interacts with applications and faces uncertainty. The included Gaia2 benchmark offers 800 scenarios across multiple “universes”. It can test reasoning, memory, tool use, and adaptability. Integration with simulated applications/agent APIs (email, file system, etc.). ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 12
    AlphaGenome

    AlphaGenome

    Programmatic access to the AlphaGenome model

    ...The model analyzes DNA sequences of up to 1 million base pairs in length and can deliver predictions at single-base-pair resolution for most outputs. AlphaGenome achieves state-of-the-art performance across a range of genomic prediction benchmarks, including numerous diverse variant effect prediction tasks.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 13
    stdlib

    stdlib

    Standard library for JavaScript and Node.js

    ...Functions to assert, group, filter, map, pluck, and transform your data both in browsers and on the server. Everything you would expect from a modern standard library. Consistent interfaces combined with extensive documentation, examples, tests, and benchmarks. High-quality implementations so you can focus less on finding the right package and more on getting work done.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 14
    GopherLua

    GopherLua

    GopherLua: VM and compiler for Lua in Go

    GopherLua is a Lua5.1(+ goto statement in Lua5.2) VM and compiler written in Go. GopherLua has the same goal as Lua: To be a scripting language with extensible semantics. It provides Go APIs that allow you to easily embed a scripting language to your Go host programs. The stack-based API like the one used in the original Lua implementation will cause a performance improvement in GopherLua (It will reduce memory allocations and concrete type <-> interface conversions). GopherLua API is not a...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 15
    Async MySQL Connector .NET and .NET Core

    Async MySQL Connector .NET and .NET Core

    Async MySQL Connector for .NET and .NET Core

    ...This library implements true asynchronous I/O for database operations, without blocking (or using Task.Run to run synchronous methods on a background thread). This greatly improves the throughput of a web server that performs database operations. This library outperforms MySQL Connector/NET (MySql.Data) on benchmarks. This library is MIT-licensed and may be freely distributed with commercial software. Commercial software that uses Connector/NET may have to purchase a commercial license from Oracle. Fixes dozens of open bugs in Oracle’s Connector/NET; passes all ADO.NET Specification Tests. First MySQL library to support .NET Core; uses the latest .NET features.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 16
    Qwen3 Embedding

    Qwen3 Embedding

    Designed for text embedding and ranking tasks

    ...It builds upon the Qwen3 base/dense models and offers several sizes (0.6B, 4B, 8B parameters), for both embedding and reranking, with high multilingual capability, long‐context understanding, and reasoning. It achieves state-of-the-art performance on benchmarks like MTEB (Multilingual Text Embedding Benchmark) and supports instruction-aware embedding (i.e. embedding task instructions along with queries) and flexible embedding/vector dimension definitions. It is meant for tasks such as text retrieval, classification, clustering, bitext mining, and code retrieval.
    Downloads: 2 This Week
    Last Update:
    See Project
  • 17
    AIDE ML

    AIDE ML

    AI-Driven Exploration in the Space of Code

    ...The project implements the AIDE algorithm, which uses a tree-search strategy guided by large language models to iteratively generate, evaluate, and refine code. Instead of relying on manual experimentation, the agent autonomously drafts machine learning pipelines, debugs errors, and benchmarks performance against user-defined evaluation metrics. The system repeatedly improves its generated code by exploring different implementation paths and selecting the best-performing solutions. AIDE ML is packaged as a Python toolkit with built-in utilities such as command-line tools, configuration presets, and visualization interfaces that allow researchers to observe how the search process evolves. ...
    Downloads: 1 This Week
    Last Update:
    See Project
  • 18
    LLM Colosseum

    LLM Colosseum

    Benchmark LLMs by fighting in Street Fighter 3

    LLM-Colosseum is an experimental benchmarking framework designed to evaluate the capabilities of large language models through gameplay interactions rather than traditional text-based benchmarks. The system places language models inside the environment of the classic video game Street Fighter III, where they must interpret the game state and decide which actions to perform during combat. This setup creates a dynamic environment that tests reasoning, situational awareness, and decision-making abilities in real time. Instead of relying purely on reward signals as in reinforcement learning agents, the models analyze contextual information and generate strategic actions based on the game environment. ...
    Downloads: 1 This Week
    Last Update:
    See Project
  • 19
    Tencent-Hunyuan-Large

    Tencent-Hunyuan-Large

    Open-source large language model family from Tencent Hunyuan

    Tencent-Hunyuan-Large is the flagship open-source large language model family from Tencent Hunyuan, offering both pre-trained and instruct (fine-tuned) variants. It is designed with long-context capabilities, quantization support, and high performance on benchmarks across general reasoning, mathematics, language understanding, and Chinese / multilingual tasks. It aims to provide competitive capability with efficient deployment and inference. FP8 quantization support to reduce memory usage (~50%) while maintaining precision. High benchmarking performance on tasks like MMLU, MATH, CMMLU, C-Eval, etc.
    Downloads: 1 This Week
    Last Update:
    See Project
  • 20
    Perf Book

    Perf Book

    The book "Performance Analysis and Tuning on Modern CPU"

    ...It explains how caches, TLBs, prefetchers, branch predictors, and out-of-order execution influence real program speed, then connects those concepts to concrete optimization strategies. Readers learn how to design trustworthy benchmarks, avoid measurement traps (warmup, turbo, frequency scaling), and interpret hardware performance counters. The book walks through vectorization, memory layout, data-oriented design, and algorithmic choices, illustrating when compiler flags, intrinsics, or hand-rolled assembly make sense. It also demonstrates tool-driven workflows—using profilers and PMU events—to locate true bottlenecks and validate that changes actually help. ...
    Downloads: 1 This Week
    Last Update:
    See Project
  • 21
    OmniParser

    OmniParser

    A simple screen parsing tool towards pure vision based GUI agent

    ...Additionally, a collection of 7,000 icon-description pairs is used to fine-tune a caption model that extracts the functional semantics of detected elements. Evaluations on benchmarks such as SeeClick, Mind2Web, and AITW demonstrate that OmniParser outperforms GPT-4V baselines, even when using only screenshot inputs without additional information.
    Downloads: 1 This Week
    Last Update:
    See Project
  • 22
    Unity ML-Agents Toolkit

    Unity ML-Agents Toolkit

    Unity machine learning agents toolkit

    ...Using ML-Agents allows developers to create more compelling gameplay and an enhanced game experience. Advancement of artificial intelligence (AI) research depends on figuring out tough problems in existing environments using current benchmarks for training AI models. Using Unity and the ML-Agents toolkit, you can create AI environments that are physically, visually, and cognitively rich.
    Downloads: 5 This Week
    Last Update:
    See Project
  • 23
    ADR

    ADR

    ADR secures enterprise AI agents through observability

    ADR, short for Agentic AI Detection and Response, is an enterprise security system for monitoring and evaluating AI agents. It captures agent intent, tool activity, and execution traces from coding assistants, internal automations, and customer-facing agents. A normalized sensor layer provides observability across multiple agent tools and operating systems. ADR-Bench supplies more than 300 realistic tasks, 133 MCP servers, and coverage of 17 documented agent attack techniques. Its two-tier...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 24
    BaseRT

    BaseRT

    Fastest LLM inference runtime for Apple Silicon

    ...Stable C interfaces connect the engine with Python, Node.js, Rust, and Swift applications. The repository contains the open CLI, format specifications, bindings, documentation, and benchmarks, while the prebuilt inference engine uses a separate license.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 25
    Codex Autoresearch

    Codex Autoresearch

    A codex plugin for running optimization loops inside a codebase

    Codex Autoresearch is an autonomous software improvement framework that enables AI coding agents to iteratively enhance codebases without continuous human input. The system operates in a loop where the agent modifies code, evaluates results against measurable metrics, and either keeps or discards changes based on performance. It generalizes the concept of autoresearch beyond machine learning, allowing optimization of test coverage, latency, lint errors, and overall code quality. Developers...
    Downloads: 0 This Week
    Last Update:
    See Project