A microbenchmark support library
Go web framework benchmark
A command-line benchmarking tool
RandomX, KawPow, CryptoNight, AstroBWT and GhostRider unified miner
GPU benchmark testing graphics performance with realistic 3D scenes.
A benchmark built to evaluate and improve agent capabilities
A benchmarking framework for the Julia language
A Heterogeneous Benchmark for Information Retrieval
A Comprehensive Benchmark to Evaluate LLMs as Agents (ICLR'24)
A simple disk benchmark software
AI coding agent optimized for small LLMs. 87% benchmark
Agentic, Reasoning, and Coding (ARC) foundation models
Checks whether Kubernetes is deployed
Meta Agents Research Environments is a comprehensive platform
LongBench v2 and LongBench (ACL 25'&24')
A.S.E (AICGSecEval) is a repository-level AI-generated code security
Drill is an HTTP load testing application written in Rust
Autonomous red teaming platform
CrystalMark Retro is a comprehensive benchmarking software.
MTEB: Massive Text Embedding Benchmark
AI framework to autonomously improve the performance of any AI system
A high-performance HTTP benchmarking tool
Benchmarking synthetic data generation methods
Code for the paper "Evaluating Large Language Models Trained on Code"
Integrates the JMH benchmarking framework with Gradle