scout.

a daily read of the ML and AI papers

TUE · 29 SEP 2026
3 papers

Billed twice for the same read

Three papers on what you hand a coding agent and what it does with it: the waste sitting in its own trajectories, the documentation it turns out not to need, and the effects nobody approved.

Today's pick
up to 41.73%
lower cost per task from seven lines of hand-written guidance, against about 20.8% from skills the agent wrote for itself

Analyzing and Mitigating Cost-Inefficient Behaviors in Coding Agents

A coding agent re-reads code a subagent already fetched, writes a fourth near-copy of the same throwaway test script, and runs an unchanged test suite eight times — and you pay for all of it. Across 1,200 trajectories from two agents on 300 real repository-issue tasks, those three behaviours appear in 79.00% to 98.00% of tasks and account for 6.86% to 22.75% of task cost, worst in the cheap-model configurations. Two obvious fixes disappoint: bolting on graph-based structural retrieval raised cost by as much as 28.14% while shifting how the agent delegates, and skills the agent synthesised from its own traces came out low-level and trace-specific, saving about 20.8% at best. What worked was seven high-level, trace-agnostic instructions a developer wrote once — state a hypothesis before retrieving, revise a script instead of cloning it, rerun tests only when the code changed — cutting cost by up to 41.73% with pass rates within a few points either way.

0.08 → 0.71
share of tests passed when the agent cannot see the source and gets a well-written description instead — the same descriptions gain nothing once it can

Compact Documentation for Coding Agents: A Benchmark, an Optimizer, and Why It Does Not Transfer

Writing a compact summary of every module so your coding agent has "context" feels obviously right, and this team built the tooling to do it properly: a benchmark that scores a description by whether code regenerated from it alone still passes the original tests, and a prompt-search loop that reached full fidelity on files it had never seen. Those descriptions are powerful when the code is out of reach — mean share of tests passed rises from 0.08 on the issue alone to 0.71 — and worth nothing when it is not. Across two model families, ten repositories and 99 tasks, no documentation condition beat the issue by more than a single resolved task and the long full-length description was the worst arm wherever it ran, while the benchmark's own retrieved context did produce a real gain (15 of 29 against 14 of 29). A description faithful enough to rebuild a file is a restatement of that file: redundant while the file is in context, and occasionally harmful because it nudges the agent to rewrite rather than patch. Spend tokens on docs only where the source cannot fit, and buy completeness, not length.

+1.08 ms
median added time to check that what persisted is what the application approved — 618 of 618 clean runs preserved, every injected wrong commit blocked

Beyond Approved Actions: Runtime Validation of Persistent Outcomes in Agent Workflows

Your agent's tool call returns success, and the row it was approved to update really did change — along with a notification a database trigger inserted, or a second charge because the first response was lost and the call was retried. Approving the call and logging the aftermath both miss this: nothing compares what actually persisted against what was authorised before the next step consumes that success. The runtime here stages writes inside a private transaction, collects the effects that would persist, and commits only on an exact match against the approved outcome for this state, plan and single execution: on 206 public business tasks it kept all 618 correct outcomes, rejected all 354 omission and 138 duplication variants that an error-reject workflow accepted, and blocked every invalid handoff across 80 dependency topologies, for about 1.08 ms of added time at the median. If your agent writes into systems with triggers, webhooks or retries, the cheap version is to read the effect set back after the write and treat anything you did not authorise as a failed step rather than a logged anomaly.