Harvey LAB
A benchmark built to evaluate and improve agent capabilities
...The environment is designed to test practical legal workflows rather than isolated question-answering ability. Evaluation tools score agent outputs and support reports and comparative experiment runs. Documentation includes an end-to-end M&A data-room example covering setup, execution, scoring, and result analysis. The project is intended to help researchers and developers identify weaknesses and measure improvements in legal AI agents.