/p/2026-10-12 · explainer
Paper explainer · 2610.11559 · Deng, Wang, Jiang and 13 others

Your benchmark measured the expert.

Hold the model, the repository and the task fixed, and change only who is typing. Across 13 models on long multi-turn coding tasks, the tests for the requested feature passed 78.5% of the time when the simulated user was a software architect and 23.0% of the time when it was a non-coder — a gap of 55.5 points that belongs entirely to the person on the other end. Most of it is localisation: with a vague user the models were looking in the wrong place on 72.2% of turns.

01 · The problem

The same agents, four different users

who is talking to the agent

02 · The mechanism

Where the four users came from

pick a persona

step through what the simulated user does each turn

03 · Where it goes wrong

It is not the fixing, it is the finding

which failure you are looking at
and who was talking

04 · Why I care

Talking to a vague user is worse than not talking

model

05 · Apply it

What your own number is worth illustrative

the persona your eval actually uses

Results

What the paper actually measured

What it does not show

In practice