Showing 728 open source projects for "benchmark"

View related business solutions
  • MongoDB Atlas runs apps anywhere Icon
    MongoDB Atlas runs apps anywhere

    Deploy in 115+ regions with the modern database for every enterprise.

    MongoDB Atlas gives you the freedom to build and run modern applications anywhere—across AWS, Azure, and Google Cloud. With global availability in over 115 regions, Atlas lets you deploy close to your users, meet compliance needs, and scale with confidence across any geography.
    Start Free
  • Veeam Data Platform v13.1 Icon
    Veeam Data Platform v13.1

    Move workloads across hypervisors and clouds with no vendor lock-in. Try VDP free today.

    Try Veeam Data Platform today. Experience the unified platform that's secure by design, portable by default, and proven to recover clean, fast, and anywhere.
    Try Now
  • 1
    DeepSeek-OCR 2

    DeepSeek-OCR 2

    Visual Causal Flow

    ...It is designed to handle complex layouts and noisy documents by giving the model causal reasoning capabilities that mimic human visual scanning behavior, enhancing OCR performance on documents with rich spatial structure. The repository provides model code and inference scripts that let researchers and developers run and benchmark the system on both images and PDFs, with support for batch evaluation and optimized pipelines leveraging vLLM and transformers.
    Downloads: 3 This Week
    Last Update:
    See Project
  • 2
    Hallucination Leaderboard

    Hallucination Leaderboard

    Leaderboard Comparing LLM Performance at Producing Hallucinations

    ...By focusing on hallucination rates rather than traditional metrics such as accuracy or fluency, the benchmark highlights an important aspect of AI system safety and trustworthiness. The leaderboard is regularly updated as new models are released and evaluation methods evolve.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 3
    SAM 3

    SAM 3

    Code for running inference and finetuning with SAM 3 model

    ...This capability is grounded in a new data engine that automatically annotated over four million unique concepts, producing a massive open-vocabulary segmentation dataset and enabling the model to achieve 75–80% of human performance on the SA-CO benchmark, which itself spans 270K unique concepts.
    Downloads: 11 This Week
    Last Update:
    See Project
  • 4
    golang-set

    golang-set

    A simple generic set type for the Go language

    ...One common interface to both implementations, a nonthreadsafe implementation favoring performance, a threadsafe implementation favoring concurrent use. Feature complete set implementation modeled after Python's set implementation. Exhaustive unit-test and benchmark suite. This package is trusted by many companies and thousands of open-source packages. This package now fully supports generic syntax so you are now able to instantiate a collection for any comparable type object.
    Downloads: 1 This Week
    Last Update:
    See Project
  • Paessler: Easy to Use With Enterprise Power. Free Trial Icon
    Paessler: Easy to Use With Enterprise Power. Free Trial

    A low-code dashboard makes monitoring intuitive for any admin, while scripting and custom sensors give experts full control.

    You shouldn't have to choose between a monitoring tool that's easy to use and one that's powerful enough for a complex environment. PRTG's low-code interface lets any admin build dashboards, set alerts and monitor devices without scripting, while custom sensors and full API access are there when your team needs deeper control. One platform, no compromise. Download a free 30-day trial now.
    Get Free Download
  • 5
    Apodex FrontierAgent

    Apodex FrontierAgent

    FrontierAgent, our agent framework, open-sourced alongside it

    ...A live task board tracks pending, active, completed, blocked, and canceled work. Users can provide new instructions while an agent is running without discarding the active session. The same modular workflow engine also powers benchmark evaluations and can connect to OpenAI-compatible model endpoints.
    Downloads: 2 This Week
    Last Update:
    See Project
  • 6
    KonaBess

    KonaBess

    A GPU overclock & undervolt tool for various Snapdragon chips

    ...The application achieves customization by unpacking the Boot/Vendor Boot image, decompiling and editing relevant dtb (device tree binary) files, and finally repacking and flashing the modified image. The extent of improvement varies, with some users reporting a 25% reduction in power consumption in the graphics benchmark (4.2w->3.2w) after undervolting the Snapdragon 865. Actual improvement is chip-specific and contingent on stability requirements.
    Downloads: 28 This Week
    Last Update:
    See Project
  • 7
    Napkin Math

    Napkin Math

    Techniques and numbers for estimating system's performance

    Napkin Math is a technical reference project for estimating software system performance from first principles. It collects practical numbers, benchmark-style measurements, and mental models that help engineers make fast back-of-the-envelope calculations. The project is useful for questions like how much memory throughput matters, how long storage operations may take, what network latency to expect, or how expensive logging could become at high request volume. It treats these values as rounded numbers for reasoning rather than exact performance guarantees. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 8
    Open Source Vizier

    Open Source Vizier

    Python-based research interface for blackbox

    ...Defines abstractions and utilities for implementing new optimization algorithms for research and to be hosted in the service. A wide collection of objective functions and methods to benchmark and compare algorithms. Define a problem statement and study configuration. Setup a local server, setup a client to connect to the server, perform a typical tuning loop, and use other client APIs.
    Downloads: 1 This Week
    Last Update:
    See Project
  • 9
    LLM Colosseum

    LLM Colosseum

    Benchmark LLMs by fighting in Street Fighter 3

    LLM-Colosseum is an experimental benchmarking framework designed to evaluate the capabilities of large language models through gameplay interactions rather than traditional text-based benchmarks. The system places language models inside the environment of the classic video game Street Fighter III, where they must interpret the game state and decide which actions to perform during combat. This setup creates a dynamic environment that tests reasoning, situational awareness, and decision-making...
    Downloads: 1 This Week
    Last Update:
    See Project
  • Host LLMs in Production With On-Demand GPUs Icon
    Host LLMs in Production With On-Demand GPUs

    NVIDIA L4 GPUs. 5-second cold starts. Scale to zero when idle.

    Deploy your model, get an endpoint, pay only for compute time. No GPU provisioning or infrastructure management required.
    Start Free
  • 10
    NYC Taxi Data

    NYC Taxi Data

    Import public NYC taxi and for-hire vehicle (Uber, Lyft)

    ...It also contains example analyses—spatial and temporal visualizations like maps, time-series plots, and hotspot detection—highlighting insights such as patterns of demand, peak times, and geospatial distributions. The repository is often used as a benchmark dataset and example for teaching, benchmarking, and demonstration purposes in the data science and urban analytics communities.
    Downloads: 1 This Week
    Last Update:
    See Project
  • 11
    AutoAgent AI

    AutoAgent AI

    Autonomous harness engineering

    ...Instead of manually tuning prompts or workflows, developers define high-level goals in a configuration file, and the system continuously modifies its own tools, orchestration, and logic based on benchmark performance. It operates through a loop of testing, analyzing failures, and refining the agent’s configuration to maximize a scoring metric. The framework uses a single-file agent harness combined with structured tasks and evaluation suites to guide optimization. It runs inside Docker for safe execution and reproducibility. This approach shifts agent development from manual design to automated optimization. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 12
    Recursive Language Models

    Recursive Language Models

    General plug-and-play inference library for Recursive Language Models

    ...Within the framework, you can define custom agents, environments, policy networks, and reward structures while leveraging built-in dataset utilities, logging, and checkpointing for reproducible experiments. RLM also includes integration with popular simulation environments and benchmark suites, giving researchers a ready-made playground for algorithm comparison and performance tracking.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 13
    D4RL

    D4RL

    Collection of reference environments, offline reinforcement learning

    D4RL (Datasets for Deep Data-Driven Reinforcement Learning) is a benchmark suite focused on offline reinforcement learning — i.e., learning policies from fixed datasets rather than via online interaction with the environment. It contains standardized environments, tasks and datasets (observations, actions, rewards, terminals) aimed at enabling reproducible research in offline RL. Researchers can load a dataset for a given task (e.g., maze navigation, manipulation) and apply their algorithm without the need to collect fresh transitions, which accelerates experimentation and comparison. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 14
    Behaviour Suite Reinforcement Learning

    Behaviour Suite Reinforcement Learning

    bsuite is a collection of carefully-designed experiments

    ...Its main goal is to identify, measure, and analyze fundamental aspects of learning efficiency and generalization in RL algorithms. The library enables researchers to benchmark their agents on standardized tasks, facilitating reproducible and transparent comparisons across different approaches. Each experiment in bsuite is meticulously designed to capture key challenges in RL, such as exploration, credit assignment, and stability. The framework supports automated logging and analysis, generating standardized output compatible with Jupyter notebooks for streamlined evaluation. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 15
    undici

    undici

    An HTTP/1.1 client, written from scratch for Node.js

    An HTTP/1.1 client, written from scratch for Node.js. This section documents our most commonly used API methods. Additional APIs are documented in their own files within the docs folder and are accessible via the navigation list on the left side of the docs site. Garbage collection in Node is less aggressive and deterministic (due to the lack of clear idle periods that browsers have through the rendering refresh rate) which means that leaving the release of connection resources to the...
    Downloads: 15 This Week
    Last Update:
    See Project
  • 16
    BaseRT

    BaseRT

    Fastest LLM inference runtime for Apple Silicon

    ...It accelerates model execution through hand-written Metal kernels and requires an M1 or newer Mac running macOS 14 or later. A unified command-line interface can download models from Hugging Face, convert checkpoints, launch chats, benchmark performance, and inspect model packages. Its server implements OpenAI-compatible chat, completion, embedding, transcription, tool-calling, and multimodal endpoints. The custom .base format supports affine quantization from Q2 through Q8, optional AWQ calibration, and signed model bundles. Stable C interfaces connect the engine with Python, Node.js, Rust, and Swift applications. ...
    Downloads: 35 This Week
    Last Update:
    See Project
  • 17
    pnpm

    pnpm

    Fast, disk space efficient package manager

    Fast, disk space efficient package manager. pnpm uses a content-addressable filesystem to store all files from all module directories on a disk. When using npm, if you have 100 projects using lodash, you will have 100 copies of lodash on disk. With pnpm, lodash will be stored in a content-addressable storage. Files inside node_modules are cloned or hard-linked from a single content-addressable storage. pnpm has built-in support for multiple packages in a repository. pnpm creates a non-flat...
    Downloads: 17 This Week
    Last Update:
    See Project
  • 18
    Mitata

    Mitata

    benchmark tooling that loves you

    A high-performance JavaScript benchmarking tool designed for fast and accurate performance testing.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 19
    CUDA Agent

    CUDA Agent

    Large-Scale Agentic RL for High-Performance CUDA Kernel Generation

    ...The system operates in a ReAct-style loop where the agent profiles baseline implementations, writes CUDA code, compiles it in a sandbox, and iteratively refines performance. CUDA-Agent has demonstrated strong benchmark results, achieving high pass rates and significant speedups compared with compiler baselines such as torch.compile.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 20
    Image Harmonization Dataset iHarmony4

    Image Harmonization Dataset iHarmony4

    The first large-scale public benchmark dataset for image harmonization

    ...The iHarmony4 dataset comprises four sub-datasets (HCOCO, HAdobe5k, HFlickr, Hday2night), each making composite images by combining a foreground from one image with a background from another, along with associated ground truth harmonized images and foreground masks. The dataset is intended as a benchmark resource to enable and standardize research in image harmonization. Each composite sample has: composite image, foreground mask, and corresponding real harmonized image.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 21
    ReinforcementLearning.jl

    ReinforcementLearning.jl

    A reinforcement learning package for Julia

    A collection of tools for doing reinforcement learning research in Julia. Provide elaborately designed components and interfaces to help users implement new algorithms. Make it easy for new users to run benchmark experiments, compare different algorithms, and evaluate and diagnose agents. Facilitate reproducibility from traditional tabular methods to modern deep reinforcement learning algorithms. Make it easy for new users to run benchmark experiments, compare different algorithms, and evaluate and diagnose agents. Facilitate reproducibility from traditional tabular methods to modern deep reinforcement learning algorithms. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 22
    PyTorch Geometric

    PyTorch Geometric

    Geometric deep learning extension library for PyTorch

    It consists of various methods for deep learning on graphs and other irregular structures, also known as geometric deep learning, from a variety of published papers. In addition, it consists of an easy-to-use mini-batch loader for many small and single giant graphs, a large number of common benchmark datasets (based on simple interfaces to create your own), and helpful transforms, both for learning on arbitrary graphs as well as on 3D meshes or point clouds. We have outsourced a lot of functionality of PyTorch Geometric to other packages, which needs to be additionally installed. These packages come with their own CPU and GPU kernel implementations based on C++/CUDA extensions. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 23
    Rbatis

    Rbatis

    Rust High Performance compile-time ORM(RBSON based)

    ...Zero cost Dynamic SQL, implemented using (proc-macro,compile-time, Cow (Reduce unnecessary cloning)) techniques, don't need ONGL engine(mybatis) Free deserialization, Auto Deserialize to any struct(Option,Map,Vec...) High performance, Based on Future, with async_std/tokio, single threaded benchmark can easily achieve 200,000 QPS. logical deletes, pagination, py-like SQL and basic Mybatis functionalities. Supports logging, customizable logging based on log crate. 100% Safe Rust with #![forbid(unsafe_code)] enabled. rbatis/example (import into Clion!). abs_admin project an complete background user management system( Vue.js+rbatis+actix-web).
    Downloads: 14 This Week
    Last Update:
    See Project
  • 24
    msgspec

    msgspec

    A fast serialization and validation library, with builtin

    msgspec is a fast serialization and validation library, with builtin support for JSON, MessagePack, YAML, and TOML.
    Downloads: 1 This Week
    Last Update:
    See Project
  • 25
    Shannon

    Shannon

    Fully autonomous AI hacker to find actual exploits in your web apps

    Shannon is an autonomous AI penetration testing system built to find and prove real, exploitable vulnerabilities in web applications rather than stopping at static warnings or best-guess alerts. It focuses on “proof by exploitation,” meaning it actively hunts for attack vectors in your code and then attempts to execute end-to-end exploits to demonstrate impact. The project blends source-aware analysis with automated web interaction so it can validate issues like injection flaws,...
    Downloads: 27 This Week
    Last Update:
    See Project