A Statistical Acceptance Gate for Self-Evolving Agents
An agent that improves itself by rewriting its own skill file — the workflow notes, tool rules and decision logic it carries between runs — needs a rule for which proposed edit to keep, and the usual one keeps anything that lifts the average validation score. That ships regressions quietly, because an edit can raise the mean while breaking cases the old skill already handled, and it is fooled by luck: the best score on a small noisy validation set is biased upward, so edits get committed for winning a coin toss. Stop comparing averages instead — run both versions on the identical items, count only the items where one wins and the other loses, weight a broken item more heavily than a newly solved one, and commit only when a one-sided test on those counts beats chance. Across five tasks and four models at matched budget that cut already-solved items broken from 36.5% to 0% on a maths set and 42.8% to 0% on an office-document set, and scored highest in all twenty settings; the item-by-item comparison alone is worth 6.20 of the 8.73 average points gained.