An agent that edits its own skill file has two moving parts: something that proposes the edit, and something that decides whether to keep it. Almost all the attention has gone to the proposer. The decider is usually one line — keep it if the average validation score went up — and that line is where the damage happens, because an edit can lift the average while breaking items the skill already solved, and because the best score on a small noisy set is flattering by construction. Compare the two versions item by item instead, weight a broken item above a newly solved one, and commit only when a one-sided test says the wins beat the losses beyond chance: the share of solved items broken falls from 36.5% to 0% on one task and 42.8% to 0% on another, and the final score goes up, in all twenty settings.
A self-evolving agent keeps a persistent skill document: the workflow it follows, the rules it applies to its tools, the decision logic it has accumulated. After a run it proposes an edit to that document, and something has to accept or reject it. The standard rule accepts anything that raises the aggregate score on a validation set.
That rule cannot distinguish between two very different edits with the same average. Move the sliders: hold the net gain steady and change how it was earned.
The aggregate gate sees one number and it is the same number in both cases. The paired view sees two, and the second one is irreversible: a committed skill is what the agent carries into every later run, so an item it used to solve and now fails is not a dip in a metric, it is a capability deleted.
Here are three real candidate edits as the two gates score them. Only the middle one is a genuine improvement, and only one gate can tell.
And this is what the difference costs in practice. Each row is one task; the open dot is the share of already-solved items that accepted edits broke under the aggregate gate, and the filled dot is the same share under the paired gate.
The first half of the fix is bookkeeping. Run the current skill and the edited skill on the same validation items and record, per item, which one got it right. Items where both agree tell you nothing about the edit and are discarded. What is left is a ledger of wins and losses.
The second half is a value judgement, and it is the one worth arguing about in your own system: a loss is worth more than a win, because a loss is permanent and a missed win is not. The weight λ says how much more. At λ = 1 you are back to counting items. Push it up and the same ledger stops clearing.
There is a third dial, a hard cap on the share of already-solved items an edit may break at all. It is the part you can ship this afternoon without any statistics: whatever else an edit does, refuse it if it breaks more than τ of what already works.
Bookkeeping alone still commits edits that got lucky. If the proposer generates several candidates and you keep the best-scoring one, the score you kept is biased upward: you selected on noise as well as on merit, so the winner’s true effect is smaller than its measured one. That gap is not a rounding error, and it grows with the number of candidates you choose between.
So the gate asks a question with a right answer: given w wins and ℓ losses on paired items, how surprising would that split be if the edit were no better than what it replaces? Under the loss weight λ, “no better” means a win rate of λ / (1 + λ) among the items that moved, and the answer is the upper tail of a binomial. Commit when that tail falls below α. Otherwise abstain — keep the skill you have and spend the next proposal.
Two properties are worth noticing before you copy this. At λ = 1, α = 1 and no cap, the gate reduces exactly to the rule it replaces — it is a strict tightening, not a different policy, so you can dial it back to your current behaviour and step forward from there. And it is genuinely conservative: four wins with no losses at all still fails at α = 0.05, because four coin flips going your way is not evidence. Small consistent improvements will be abstained on. That is the price.
The counter-intuitive result is that committing fewer edits leaves the agent better off. Same proposer, same number of proposals, same token budget — only the gate differs. Step through the tasks.
The ablation says where the gain lives, and it is mostly in the bookkeeping rather than the statistics. Scoring per item and penalising breaks, with no test at all, is worth 6.20 points on average across the twenty settings; adding the test adds 2.53 more. If you only do one of the two, do the ledger.
You probably do not have an agent rewriting its own skill file yet. You almost certainly have a prompt, a tool description or a router threshold that somebody changes weekly, and an eval set whose pass rate decides whether the change ships. That is the identical situation: a proposer, a noisy finite validation set, and a gate made of one comparison.
Set the size of your eval set, what your current version passes, and what one proposed change did to the items. The verdicts are computed with the paper’s test; the volume is yours.
The cheap version of all of this is one column in your eval report: not the pass rate, but the list of items that passed last week and fail this week. Nobody argues with that list, and the aggregate score will never show it to you.