scout.

a daily read of the ML and AI papers

THU · 01 OCT 2026
3 papers

It scored better and broke what worked

Three papers on changes that look like improvements: the agent edit that raises your average, the merged pull request that claims a speed-up, and the second pass that overturns a right answer.

Today's pick
36.5% → 0%
of already-solved tasks broken by accepted edits to an agent's skill file, once the gate compares both versions item by item instead of on the average

A Statistical Acceptance Gate for Self-Evolving Agents

An agent that improves itself by rewriting its own skill file — the workflow notes, tool rules and decision logic it carries between runs — needs a rule for which proposed edit to keep, and the usual one keeps anything that lifts the average validation score. That ships regressions quietly, because an edit can raise the mean while breaking cases the old skill already handled, and it is fooled by luck: the best score on a small noisy validation set is biased upward, so edits get committed for winning a coin toss. Stop comparing averages instead — run both versions on the identical items, count only the items where one wins and the other loses, weight a broken item more heavily than a newly solved one, and commit only when a one-sided test on those counts beats chance. Across five tasks and four models at matched budget that cut already-solved items broken from 36.5% to 0% on a maths set and 42.8% to 0% on an office-document set, and scored highest in all twenty settings; the item-by-item comparison alone is worth 6.20 of the 8.73 average points gained.

18 of 30
merged agent speed-up fixes that delivered at least half the gain they claimed, when re-run on three workload sizes

Merged, Not Measured: Performance Issues Fixed by Coding Agents

Coding agents open pull requests claiming a speed-up, and maintainers merge 57% of the ones they close — but a merge is a social signal, not a measurement. Re-running 30 merged performance fixes, 18 delivered at least half the claimed gain, 3 fell short, 9 showed no significant gain or ran slower, and 14 changed behaviour on inputs no test covered; only 11% of the 1,262 fixes studied carried a performance test, and carrying one did not raise the odds of a merge. Acceptance tracked history rather than content: 31–37% for an agent new to a repository against 70% once it had a record there, and 33% up to 84% as the repository's own prior merge rate on agent pull requests rose. Make the benchmark part of the pull request and re-run it yourself, because nothing in the review path you have now is checking the claim.

19.1%
of initially correct answers turned wrong by a second self-review pass, on the same model that gained 25.5 points overall

When Should LLMs Trust Their Own Revisions?

Asking a model to look again at its own answer, with no new evidence, runs two effects at once: it repairs some wrong answers and overturns some right ones, and the accuracy figure you read is only their difference. Over 29 open-weight models on a yes/no reading task, grade-school maths and a correlation-versus-cause set, the two come apart — one 8-billion model gained 25.5 points on the maths set while turning 19.1% of its initially correct answers wrong, and on the reading task a blanket second pass lowered accuracy in 20 of 28 settings. All five revision prompts tested lost accuracy on average, and the harder the prompt pushed the model to reconsider, the more right answers it broke. Score self-correction with two numbers rather than one, and gate it on a cheap first-pass signal such as log-probability spread; gating won in 50 of 82 settings and saved a median 56.5% of revision tokens on the reading task.