CONTRA: Discovering and Qualifying Behavior-Changing Questions for Selective Clarification in LLM Code Generation
A coding agent that meets an underspecified requirement does not stop and ask; it quietly resolves the gap with its own assumption, and the cost of undoing that grows as later work builds on it. CONTRA decides which questions are worth interrupting you for by testing them rather than judging them: for each candidate it writes the two plausible answers out as rival versions of the requirement, generates three programs per answer, runs both groups on at least seven shared inputs, and keeps the question only when the two sides give different stable outputs on some input. Across four coding agents on 499 requests — 419 of them deliberately underspecified — that scored 41.2 against the best existing method's 27.3 at matching the clarifications a human had marked necessary, and on one open-weight model it roughly doubled what two popular coding harnesses ask unprompted (19.9 against 9.6 and 10.6). Copy the test rather than the framework: generate both branches and check whether the answer changes observable behaviour before putting the question in front of a user.