/p/2026-10-05 · explainer
Paper explainer · 2610.01769 · Fang, Li, Zhang, Jin et al.

Ask only when the answer changes the code.

A coding agent that hits a gap in your request does not stop — it picks a reading and keeps going, and the guess gets built on. Asking about every gap is worse: the interruptions cost more than the bugs. This paper decides which questions are worth your attention by running them: write the two plausible answers out as rival versions of the requirement, generate three programs for each, run both sets on at least seven shared inputs, and keep the question only when the two sides give a different stable answer somewhere. Across four coding agents that scored 41.2 at matching the clarifications a human had marked necessary, against 27.3 for the best existing method.

01 · The problem

Asking nothing and asking everything both score well

score it by

02 · The mechanism

Qualify a question by running it, not by reading it illustrative

03 · Which stage earns its keep

Take a stage out and watch which number falls

score it by

04 · Against the harness you already run

Two popular coding harnesses, asked the same 499 requests

score it by

05 · In your own product

What the extra run costs at your volume illustrative

Results

What the paper actually measured

What it does not show

In practice