scout.

a daily read of the ML and AI papers

MON · 05 OCT 2026
3 papers

It never asked which behaviour you meant

Three papers on the parts of agent-written code that functional tests never look at: the question it decided not to ask, the import line it invented, and the dependency list it cannot reproduce twice.

Today's pick
+13.88 pts
better at matching the clarifications a human had marked necessary (percentage points of F1), against the best existing method

CONTRA: Discovering and Qualifying Behavior-Changing Questions for Selective Clarification in LLM Code Generation

A coding agent that meets an underspecified requirement does not stop and ask; it quietly resolves the gap with its own assumption, and the cost of undoing that grows as later work builds on it. CONTRA decides which questions are worth interrupting you for by testing them rather than judging them: for each candidate it writes the two plausible answers out as rival versions of the requirement, generates three programs per answer, runs both groups on at least seven shared inputs, and keeps the question only when the two sides give different stable outputs on some input. Across four coding agents on 499 requests — 419 of them deliberately underspecified — that scored 41.2 against the best existing method's 27.3 at matching the clarifications a human had marked necessary, and on one open-weight model it roughly doubled what two popular coding harnesses ask unprompted (19.9 against 9.6 and 10.6). Copy the test rather than the framework: generate both branches and check whether the answer changes observable behaviour before putting the question in front of a user.

84%
of generated files carrying at least one library fault — wrong import path, missing import, invented package or deprecated call

Understanding and Mitigating Library-Related Issues in LLM-Generated Code

When generated code that leans on a third-party library breaks, the fault is usually in the import line rather than the logic. Hand-checking 100 generated files found 84% carried at least one library fault and 228 faults in all, about 2.7 per broken file: wrong import paths (93), missing imports (79), packages that do not exist (28), deprecated calls (17) and dead imports (11). The remedy tested here is unglamorous and easy to rebuild — have the agent name the libraries the task needs, fetch the official documentation for each, pass the verified import lines and usage snippets into generation, then compile, check the imports back against those docs and loop on failure — and on 300 tasks drawn from fast-moving agent frameworks it cut library faults by 38.1% to 54.6% across five models and lifted the share of files that compile and use their libraries correctly from 61% to 77% on the weakest. Removing the documentation fetch hurt more than removing any other stage, so if you add one step to your codegen path, add that one.

7%
overlap between two agents' dependency lists for the same task in the same language — and no two runs of one agent ever matched

Code That Works, Environments That Don't: Measuring Environment Reproducibility in AI-Generated Software

Functional tests ask whether the code runs here; they never ask whether the dependency list it shipped would rebuild it anywhere else. Three coding agents ran fifty tasks in four languages while the packages each program actually loaded were traced at the system-call level, and outside Python the declared list and the real one came apart badly: scored against the traced truth, manifest accuracy was 0.77 in Python but 0.22 in JavaScript and 0.02 in C++, where roughly half of declared packages were never loaded and 62% of what JavaScript did load was never declared. Two agents handed the same task in the same language agreed on as little as 7% of their package lists, and three runs of the same agent never once produced an identical list. Every agent eventually reached a passing run, which is the trap: a scoreboard reading only final success calls all three near-perfect, so pin the dependency list from a traced install in a clean container rather than from the manifest the agent wrote.