October 3, 2026ResearchBenchmarkAgents

Your Agent Leaderboard Is Ranking Harnesses, Not Models

Two papers this week land on the same uncomfortable conclusion: when you compare agent models, the harness can flip the result.

Finding the Right Fit ran 66 configurations: four configurable harnesses (OpenHands, DeepSeek Harness, Pi and openJiuwen) times five models on TUA-Bench, ALE-CLI and Terminal-Bench 4, plus the native Codex-GPT and Claude Code-Claude pairings. Model rankings reversed across harnesses. On Terminal-Bench 4, Claude leads GPT by 7.94 points inside OpenHands and trails it by 30.16 points inside Pi. For four of five models, the best harness changes from one benchmark to the next. A model's own vendor harness is not reliably its best, and paying more does not reliably buy a higher score: GPT scores higher under Pi than under DeepSeek Harness at less than a quarter of the cost per task. The traces explain why. Models start almost every repair themselves, so what decides the outcome is whether the harness hands the failure back in a form the model can use. GPT likes Pi's lean scaffold. Kimi, which often emits malformed tool calls, does best in openJiuwen, which wins for Kimi on all three benchmarks by 5.61 to 11.11 points.

Agent Evaluation Reliability puts numbers on the damage. Applying a Bayesian variance decomposition to 22 benchmarks from the Holistic Agent Leaderboard and Harbor Index, it finds that fixed model-plus-scaffold systems are ranked reliably (0.935 to 0.994). Underlying models are not: reliability drops to anywhere between 0.148 and 0.841. And adding more tasks of the same kind can't fully fix it.

Put together, a headline like "Model X beats Model Y on Terminal-Bench" is close to meaningless unless the harness is named. The useful unit is the pair. If you're choosing a model for your own agent, test it inside your own harness, because someone else's leaderboard measured someone else's harness. Code and data for the first paper are at github.com/liyix/finding-the-right-fit.

Links: arxiv.org/abs/2610.00917, arxiv.org/abs/2610.00651
← Previous
Supabase Buys Turso Because Agents Now Spin Up a Million Databases a Week
Next β†’
PoS: Stop Giving Agents Memory, Give Them a Belief State
← Back to all articles

Comments

Loading...
>_