October 4, 2026ResearchCodingMonitoring

Weak Reviewers Can Audit Strong Coding Agents, If You Give Them Evidence

Coding agents are good at producing patches that look done and quietly skip required behavior. Long traces and confident summaries hide the gap. Groundability, Not Scale Alone (arXiv 2610.01023) asks when a smaller, weaker model can reliably catch that. Their answer: reviewer size barely matters, evidence does.

They studied 411 execution-labeled traces from three agents plus 101 controlled cases. Giving reviewers structured but unchecked evidence on 154 GPT-5.4 traces raised defect catch, but also raised over-rejection. With official execution evidence and a frozen format, five of six reviewers improved on both rates on 122 held-out traces, and two classified every single trace correctly. Size was not a consistent predictor of quality.

Deployment doesn't have official tests, so they built a frozen cascade: patch-caused static errors, then generated tests that must first fail on the unpatched repo. On held-out GPT-5.4 and Gemini traces it reaches 0.89 and 0.86 coverage with 0.76 and 0.80 catch. But over-rejection stays at 0.66 and 0.67. It catches most bad patches and also throws out most of the good ones that reach the reviewer.

That's the honest takeaway. Cheap oversight of expensive agents is possible, and the bottleneck isn't the judge, it's producing decisive checks without a ground-truth test suite. Whoever automates "write the test that fails before the patch" well owns the review layer.

Link: arxiv.org/abs/2610.01023
← Previous
Chaining Skills Hijacks Agents 74% of the Time, and the Fix Breaks Real Work
Next β†’
Incident-Arena: Frontier Agents Fix Under 65% of Real Production Outages
← Back to all articles

Comments

Loading...
>_