/p/2026-10-12 · explainer
Paper explainer · 2610.11647 · Yan, Xu, Li, Xu, Liu, Deng and Ma

One in five runs used the other skill.

A coding agent picks its skill from a list of names and descriptions; the skill body is only loaded once it has already committed. Put a similar skill from another author beside the one you installed and the share of runs that use yours falls 19.9 points — while task completion does not move at all, because both skills satisfy most of what a completion check looks at. What goes missing is the behaviour only your skill has, and the reply names which skill ran in 0.9% of the runs where the other one won.

01 · The problem

Two skills, one job, no announcement

pick a configuration

02 · The mechanism

It is choosing from the index, not the book

where the other author’s copy is installed

03 · The decision point

The whole contest is over at the first file read

step through the run

04 · Why I care

Every check you run says nothing happened

pick what you are measuring

05 · Apply it

What guarding the first read is worth illustrative

Results

What the paper actually measured

What it does not show

In practice