scout.

a daily read of the ML and AI papers

MON · 12 OCT 2026
3 papers

One in five runs used the skill you did not install

Three papers about the scaffolding rather than the model: the skill library the agent picks from, the tests it writes to protect you, and the person on the other end. In all three the task still passes and the part you wanted quietly goes missing.

Today's pick
−19.9 pts
fewer runs used the skill you installed once a similar one from another author sat beside it (percentage points)

One Skill Too Many: How Co-Installed Skills Conflict in Coding Agents

You install a skill so the agent does a job your way, someone else's similar skill is sitting in the same project, and the agent opens theirs instead. The agent chooses between skills from names and descriptions alone — the body is only loaded after it has committed — and across 6,368 runs on three models, putting a similar skill beside yours cut the share of runs that used yours by 19.9 points, where an unrelated skill cost only 5.9. The pass signal shows none of it: task completion moved +1.9 points, because the two skills agree on most of what a completion check looks at, and what went missing was the functions only your skill provides, down 5.6 points. Where a skill sits decides the winner — a same-name copy in the personal directory removed yours from the listing entirely and cost 35.0 points, while reversing the listing order did nothing (+0.6) — the whole contest is settled at the first file read, and the final reply named which skill it had used in 0.9% of the runs where the other one won.

3.5–6.5×
how much more often a generated regression test asserted the bug was correct than caught it

On the Risks of using LLM-Generated Tests for Regression Testing

Point a test generator at the change you just merged and it will write tests that pass — including on the parts you got wrong. A team took 145 merged changes from three large Python projects, injected faults into the functions those changes touched, kept only the faults the existing suite already missed, then had a model write regression tests against the faulty code with the correct version withheld, accepting any test that passed what it was given. Between 8.4% and 16.9% of the generated tests asserted the faulty behaviour as expected, against 2.4% to 4.8% that caught it, three and a half to six and a half times more tests freezing the bug in than finding it; 82.8% to 90.9% of those were still passing and still guarding the wrong answer at the end of each project's later history, and the developers' own suite caught 13.8% to 30.3% of the faults. The fault is in the acceptance rule, so change it: a generated regression test that has only ever run against the code it was generated from has not been tested, and should have to fail against a deliberately broken build before it joins the suite.

23.0% vs 78.5%
tests passing for the requested feature when the user was a non-coder, against the same models with a software architect

SWE-Journey: Towards More Realistic Evaluation of Coding Assistants through Long-Horizon, Multi-Turn Interaction

Same assistant, same task, different person typing: across 13 models, the tests for the requested feature passed 23.0% of the time when the simulated user was a non-coder and 78.5% of the time when it was a software architect. A team clustered real coding-session transcripts into four interaction styles, built a simulated user for each with its own expertise, detail preference and verification habit, and ran long multi-turn tasks through them — roughly 20 subtasks each, reference patches averaging 10.2 files. Most of what the vague user costs is localisation: the models were looking in the wrong place on 72.2% of turns with a non-coder against 22.5% with an architect, and the weaker ones talked themselves further into the wrong direction instead of asking, since most models asked a clarifying question on under 20% of turns. If your product puts a non-expert in front of a coding agent, the figure from your own benchmark is the architect's figure; re-measure with a user who reports symptoms rather than files.