October 3, 2026ResearchAgents

ActiveSaddler: Harness Optimizers Were Training on the Wrong Problems

Automatic harness optimization is now a crowded research lane: let an optimizer rewrite an agent's prompts, tool interfaces and control logic from execution feedback, and watch the score climb. ActiveSaddler (HF 50 upvotes) points at the part everyone left fixed: which training scenarios produce that feedback in the first place.

Existing methods put a lot of work into how the harness gets updated, then feed the optimizer a scenario list set before training starts. But as the harness improves, the scenarios that would teach it the most change. The ones it used to fail are now easy; the useful failures are somewhere else. ActiveSaddler treats picking the next scenarios as a non-stationary bandit problem. It groups recurring failures into reusable failure-pattern "arms," estimates how much more progress each pattern can still yield, and balances going back to known weaknesses against exploring new scenarios to find fresh ones. Every optimization round updates both the list of failure patterns and their priority, so the curriculum evolves alongside the harness.

Same optimizer, same budget, only the curriculum changed: test Pass@1 rises 4.4 points on GAIA2 and 7.5 points on Terminal-Bench 2.0 against a fixed scenario order.

This is the third or fourth "harness as a search problem" paper in two weeks, after MILO and Mid-Harness. Each one finds a new knob that moves scores without touching weights. The lesson for practitioners is simpler than the bandit math: if you're iterating on your agent against the same 50 eval cases you started with, you're probably overfitting to problems you've already solved. Rotate in the failures you haven't seen yet.

Link: arxiv.org/abs/2610.00906, autosaddler-projectpage.github.io/activesaddler
← Previous
PoS: Stop Giving Agents Memory, Give Them a Belief State
Next β†’
GraphForge: 2,169 Trajectories Push a 27B Open Model Up 65 Points on GDPVal
← Back to all articles

Comments

Loading...
>_