A benchmark built to evaluate and improve agent capabilities
A Heterogeneous Benchmark for Information Retrieval
A Comprehensive Benchmark to Evaluate LLMs as Agents (ICLR'24)
Agentic, Reasoning, and Coding (ARC) foundation models
Meta Agents Research Environments is a comprehensive platform
LongBench v2 and LongBench (ACL 25'&24')
A.S.E (AICGSecEval) is a repository-level AI-generated code security
MTEB: Massive Text Embedding Benchmark
AI framework to autonomously improve the performance of any AI system
Benchmarking synthetic data generation methods
Code for the paper "Evaluating Large Language Models Trained on Code"
Visual Causal Flow
Leaderboard Comparing LLM Performance at Producing Hallucinations
Code for running inference and finetuning with SAM 3 model
FrontierAgent, our agent framework, open-sourced alongside it
Python-based research interface for blackbox
Designed for text embedding and ranking tasks
Benchmark LLMs by fighting in Street Fighter 3
Autonomous harness engineering
General plug-and-play inference library for Recursive Language Models
Collection of reference environments, offline reinforcement learning
bsuite is a collection of carefully-designed experiments
Large-Scale Agentic RL for High-Performance CUDA Kernel Generation
Geometric deep learning extension library for PyTorch
A fast serialization and validation library, with builtin