/p/2026-10-12 · explainer
Paper explainer · 2610.11559 · Deng, Wang, Jiang and 13 others
Your benchmark measured the expert.
Hold the model, the repository and the task fixed, and change only who is typing. Across 13 models on long multi-turn coding tasks, the tests for the requested feature passed 78.5% of the time when the simulated user was a software architect and 23.0% of the time when it was a non-coder — a gap of 55.5 points that belongs entirely to the person on the other end. Most of it is localisation: with a vague user the models were looking in the wrong place on 72.2% of turns.
Deng, Wang, Jiang and 13 others · 2026 · Tencent Hy AI Data and Beijing Zhongguancun Academy · 30 long-horizon instances · 594 subtasks · 13 models · four simulated user personas
the software architect — aligns first, reviews checkpoints, verifies systematically
the non-coder — describes symptoms, starts in the wrong place, cannot check the work