Better Harness is an open-source evaluation system for improving the workflows used by AI coding agents. Instead of reviewing only the final code change, it examines how an agent understood the task, executed work, validated results, delivered safely, and captured reusable lessons. It gathers project and session evidence while marking missing or unobserved behavior explicitly. Reports organize findings across five Agent Work Loop dimensions and connect each conclusion to visible supporting evidence. Supported gaps become prioritized recommendations with expected outcomes, repair boundaries, and acceptance checks. Repeated reports can be compared through a history view to observe workflow trends without claiming unsupported causation. Better Harness supports Claude Code, Codex, Qoder, and Cursor through host-specific plugins, commands, and report formats.
Features
- Five-dimension Agent Work Loop evaluation
- Project and session evidence collection
- Explicit handling of missing evidence
- Prioritized findings and repair plans
- Longitudinal workflow trend reports
- Claude Code, Codex, Qoder, and Cursor support