Agents Are Systems, Not Models: 54% of the Variance Is Just Rerunning
Run the exact same agent configuration twice and you'll often get a different answer. A new paper, Agents Are Systems, Not Models (arXiv 2610.01618, Cordelia Schmid among the authors), puts a number on it: about 54% of outcome variance comes from repeating the same configuration, not from changing anything. Most agent leaderboards report one run.
The setup is a coding agent that has to find and correctly operate a published specialist model on four scientific tasks. The authors vary five knobs: task information, reasoning, self-verification, time budget and backbone model. The biggest lever is what you tell the agent. Task information beats both time budget and model size, and it also cuts cost and improves calibration. The knobs interact too: more time only helps if the agent already has enough information or a strong enough model to use it.
The sharpest finding is about verification. Prompting the agent to check its answer barely changes its verification behavior. Giving it a dedicated verification tool changes that behavior substantially. Asking doesn't work and building does. That is the same lesson Robinson's OpenAI essay draws at the organizational level this week.
This is now the third paper in a week saying the harness and setup decide more than the model, after Finding the Right Fit and Mingbird. If you report an agent score without the configuration and the number of reruns, you're reporting noise with a model name on it. The benchmark and 18,000+ trajectories are released.
Link: arxiv.org/abs/2610.01618
← Back to all articles
The setup is a coding agent that has to find and correctly operate a published specialist model on four scientific tasks. The authors vary five knobs: task information, reasoning, self-verification, time budget and backbone model. The biggest lever is what you tell the agent. Task information beats both time budget and model size, and it also cuts cost and improves calibration. The knobs interact too: more time only helps if the agent already has enough information or a strong enough model to use it.
The sharpest finding is about verification. Prompting the agent to check its answer barely changes its verification behavior. Giving it a dedicated verification tool changes that behavior substantially. Asking doesn't work and building does. That is the same lesson Robinson's OpenAI essay draws at the organizational level this week.
This is now the third paper in a week saying the harness and setup decide more than the model, after Finding the Right Fit and Mingbird. If you report an agent score without the configuration and the number of reruns, you're reporting noise with a model name on it. The benchmark and 18,000+ trajectories are released.
Link: arxiv.org/abs/2610.01618
Comments