/p/2026-10-01 · explainer
Paper explainer · 2609.36043 · Wang, Xia, Liu et al.

It scored better and broke what worked.

An agent that edits its own skill file has two moving parts: something that proposes the edit, and something that decides whether to keep it. Almost all the attention has gone to the proposer. The decider is usually one line — keep it if the average validation score went up — and that line is where the damage happens, because an edit can lift the average while breaking items the skill already solved, and because the best score on a small noisy set is flattering by construction. Compare the two versions item by item instead, weight a broken item above a newly solved one, and commit only when a one-sided test says the wins beat the losses beyond chance: the share of solved items broken falls from 36.5% to 0% on one task and 42.8% to 0% on another, and the final score goes up, in all twenty settings.

01 · The problem

One number cannot tell you what the edit cost

A self-evolving agent keeps a persistent skill document: the workflow it follows, the rules it applies to its tools, the decision logic it has accumulated. After a run it proposes an edit to that document, and something has to accept or reject it. The standard rule accepts anything that raises the aggregate score on a validation set.

That rule cannot distinguish between two very different edits with the same average. Move the sliders: hold the net gain steady and change how it was earned.

One candidate edit, scored two ways

The aggregate gate sees one number and it is the same number in both cases. The paired view sees two, and the second one is irreversible: a committed skill is what the agent carries into every later run, so an item it used to solve and now fails is not a dip in a metric, it is a capability deleted.

Here are three real candidate edits as the two gates score them. Only the middle one is a genuine improvement, and only one gate can tell.

Three candidates, both gates’ verdicts

And this is what the difference costs in practice. Each row is one task; the open dot is the share of already-solved items that accepted edits broke under the aggregate gate, and the filled dot is the same share under the paired gate.

Share of already-solved items broken by accepted edits

02 · The mechanism

Score the edit on the items, and make a break cost more than a win

The first half of the fix is bookkeeping. Run the current skill and the edited skill on the same validation items and record, per item, which one got it right. Items where both agree tell you nothing about the edit and are discarded. What is left is a ledger of wins and losses.

The second half is a value judgement, and it is the one worth arguing about in your own system: a loss is worth more than a win, because a loss is permanent and a missed win is not. The weight λ says how much more. At λ = 1 you are back to counting items. Push it up and the same ledger stops clearing.

The same ledger at four loss weights
how much a broken item costs, against one new one

There is a third dial, a hard cap on the share of already-solved items an edit may break at all. It is the part you can ship this afternoon without any statistics: whatever else an edit does, refuse it if it breaks more than τ of what already works.

03 · The test

Six wins against four losses is a coin landing the way you hoped

Bookkeeping alone still commits edits that got lucky. If the proposer generates several candidates and you keep the best-scoring one, the score you kept is biased upward: you selected on noise as well as on merit, so the winner’s true effect is smaller than its measured one. That gap is not a rounding error, and it grows with the number of candidates you choose between.

Picking the best of N candidates on a noisy validation set illustrative
candidates the optimizer proposes per round
0

So the gate asks a question with a right answer: given w wins and ℓ losses on paired items, how surprising would that split be if the edit were no better than what it replaces? Under the loss weight λ, “no better” means a win rate of λ / (1 + λ) among the items that moved, and the answer is the upper tail of a binomial. Commit when that tail falls below α. Otherwise abstain — keep the skill you have and spend the next proposal.

The one-sided paired test, live

Two properties are worth noticing before you copy this. At λ = 1, α = 1 and no cap, the gate reduces exactly to the rule it replaces — it is a strict tightening, not a different policy, so you can dial it back to your current behaviour and step forward from there. And it is genuinely conservative: four wins with no losses at all still fails at α = 0.05, because four coin flips going your way is not evidence. Small consistent improvements will be abstained on. That is the price.

Reference verdicts at λ = 1

04 · What it buys

A gate that says no more often ends up scoring higher

The counter-intuitive result is that committing fewer edits leaves the agent better off. Same proposer, same number of proposals, same token budget — only the gate differs. Step through the tasks.

Regression rate and final score, per task
task

Final score, aggregate gate to paired gate

The ablation says where the gain lives, and it is mostly in the bookkeeping rather than the statistics. Scoring per item and penalising breaks, with no test at all, is worth 6.20 points on average across the twenty settings; adding the test adds 2.53 more. If you only do one of the two, do the ledger.

Where the gain comes from

05 · Your own stack

The same gate works on a prompt change and an eval set

You probably do not have an agent rewriting its own skill file yet. You almost certainly have a prompt, a tool description or a router threshold that somebody changes weekly, and an eval set whose pass rate decides whether the change ships. That is the identical situation: a proposer, a noisy finite validation set, and a gate made of one comparison.

Set the size of your eval set, what your current version passes, and what one proposed change did to the items. The verdicts are computed with the paper’s test; the volume is yours.

One prompt change, three gates illustrative
pass rate you would report
capabilities you would delete

The cheap version of all of this is one column in your eval report: not the pass rate, but the list of items that passed last week and fail this week. Nobody argues with that list, and the aggregate score will never show it to you.

Results

What the paper actually measured

What it does not show

In practice