Meta Agents Research Environments is a comprehensive platform
LongBench v2 and LongBench (ACL 25'&24')
A.S.E (AICGSecEval) is a repository-level AI-generated code security
Test-Time Reinforcement Learning
MTEB: Massive Text Embedding Benchmark
Autonomous harness engineering
Benchmark LLMs by fighting in Street Fighter 3
Optimize your code automatically with AI
Provider-agnostic, open-source evaluation infrastructure
MiniMax-M2, a model built for Max coding & agentic workflows
Generates original ARC-AGI-1-style tasks distribution-matched
Clean and efficient FP8 GEMM kernels with fine-grained scaling
Open-weight, large-scale hybrid-attention reasoning model
Spatiotemporal Signal Processing with Neural Machine Learning Models
ICLR2024 Spotlight: curation/training code, metadata, distribution
Open source codebase for Scale Agentex
Code for Cicero, an AI agent that plays the game of Diplomacy
Local AI file organization with categorization and rename suggestions
Chinese safety prompts for evaluating and improving the safety of LLMs
Beyond the Imitation Game collaborative benchmark for measuring
OpenMMLab Model Deployment Framework
High quality, fast, modular reference implementation of SSD in PyTorch
A MNIST-like fashion product database
Procedurally-Generated Game-Like Gym-Environments
8.5K high quality grade school math problems