A Comprehensive Benchmark to Evaluate LLMs as Agents (ICLR'24)
A.S.E (AICGSecEval) is a repository-level AI-generated code security
Leaderboard Comparing LLM Performance at Producing Hallucinations
Autonomous harness engineering
Generates original ARC-AGI-1-style tasks distribution-matched
Chinese safety prompts for evaluating and improving the safety of LLMs
A MNIST-like fashion product database
8.5K high quality grade school math problems