October 4, 2026ResearchBenchmarkAgents

Agents Are Systems, Not Models: 54% of the Variance Is Just Rerunning

Run the exact same agent configuration twice and you'll often get a different answer. A new paper, Agents Are Systems, Not Models (arXiv 2610.01618, Cordelia Schmid among the authors), puts a number on it: about 54% of outcome variance comes from repeating the same configuration, not from changing anything. Most agent leaderboards report one run.

The setup is a coding agent that has to find and correctly operate a published specialist model on four scientific tasks. The authors vary five knobs: task information, reasoning, self-verification, time budget and backbone model. The biggest lever is what you tell the agent. Task information beats both time budget and model size, and it also cuts cost and improves calibration. The knobs interact too: more time only helps if the agent already has enough information or a strong enough model to use it.

The sharpest finding is about verification. Prompting the agent to check its answer barely changes its verification behavior. Giving it a dedicated verification tool changes that behavior substantially. Asking doesn't work and building does. That is the same lesson Robinson's OpenAI essay draws at the organizational level this week.

This is now the third paper in a week saying the harness and setup decide more than the model, after Finding the Right Fit and Mingbird. If you report an agent score without the configuration and the number of reruns, you're reporting noise with a model name on it. The benchmark and 18,000+ trajectories are released.

Link: arxiv.org/abs/2610.01618
← Previous
Offrun: One Mac Workspace for Claude Code, Codex and Grok Build, Free
Next β†’
Mingbird: Small Models Fail Because of the Harness, Not the Model
← Back to all articles

Comments

Loading...
>_