October 5, 2026ResearchCodingSkills

RuleEvolve: Your CLAUDE.md Is Hand-Tuned, and an Evolution Loop Beats It

Every serious coding-agent setup now has a rules file, whether it's CLAUDE.md, AGENTS.md or .cursorrules. Almost all of them are written by hand, edited when something breaks and never measured. RuleEvolve, from Neil Zhenqiang Gong's group at Duke (Zhengyuan Jiang, Reachal Wang and colleagues), treats that file as something to optimize.

It's an evolutionary loop. Keep a pool of candidate rule sets. Each round, an LLM mutator writes variants of the existing candidates, a judge module scores them, and the best survive into the pool. Nothing in the agent or the model changes. Only the rules text evolves.

The authors test it across two coding-agent frameworks, four backbone LLMs and three benchmarks. RuleEvolve beats both hand-engineered rules and existing prompt-optimization baselines on functional correctness, code length and/or generation cost in tokens. The abstract doesn't give a single headline number. The claim is that it wins on whichever axis you're optimizing, across the grid.

The interesting tension is with a finding from the same week that developer-written skills cut agent cost about twice as much as agent-synthesized ones. Those two results don't actually conflict. The human prior is a better starting point than a blank page, and an evolution loop is a better editor than a human who only touches the file after a failure. The practical recipe is to seed the pool with your own CLAUDE.md and let the loop edit it, with a judge you trust. The judge is the part to get right, since a rules file tuned to a weak judge learns to please the judge.

Link: arxiv.org/abs/2610.00650
← Previous
Actions with Receipts: A Valid Citation and a Valid Trace Can Still Be a Lie
Next β†’
Super User Daily: 2026-10-05
← Back to all articles

Comments

Loading...
>_