Why the Agent That Ran a Test Shouldn't Grade ItSelf-grading AI agents report intent, not evidence. How we built a separate, skeptical judge for end-to-end test runs.Sep 11, 2026·4 min read