Coding agents repeat themselves in three specific ways: they retrieve code an earlier retrieval already covered, they write a near-copy of a script they already wrote, and they re-run tests nothing has changed. Over 1,200 trajectories on 300 real repository issues, at least one of these shows up in 79.00% to 98.00% of tasks and eats up to 22.75% of what the task costs you. Adding a structural retrieval tool did not fix it — on one configuration it cost 28.14% more. Letting the agent write its own skills from its own traces saved about 20.8% at best. Seven plain rules a developer wrote once, trace-agnostic and boring, cut cost by up to 41.73%.
Nobody audits agent trajectories for repetition, because a trajectory that ends in a passing test looks like a success. This paper read 1,200 of them and named three things that recur: subsumed retrieval — reading code fully covered by an earlier read; similar script generation — regenerating a throwaway script with minor changes instead of editing it; and test re-execution — running identical tests with no patch in between.
Pick an agent configuration and a habit. The trace is the paper’s own worked example on one Django issue; the numbers are measured over 300 tasks.
The frequencies are the part worth sitting with: the cheapest configuration in the study repeats a subsumed read 6.37 times per task. Each one is a file the agent has, in some form, already been shown.
Summed per configuration, the three habits account for 6.86% of task cost under Claude Code and 21.16% to 22.75% under the simpler scaffold, whichever model backs it. A subagent architecture that hands back summaries rather than raw files is doing real work here; a flat loop with one context is not.
Switch the habits off one at a time to see what each is worth.
Note which habit dominates where. Under Claude Code the waste is almost entirely re-reading (5.01 of 6.86 points); under the flat scaffold, script regeneration costs nearly as much as retrieval does, because a flat loop keeps the old script in context and writes a new one anyway — it fires 5.91 to 9.98 times more often there than under Claude Code.
The paper then runs about 10,000 trajectories over three mitigations on 200 held-out tasks: structure-aware retrieval (a code-graph tool the agent can query instead of reading files), agent-synthesized skills (the agent reads its own traces and writes its own guidance), and developer-designed skills (seven high-level rules a person wrote, no traces involved).
Choose an intervention and a configuration. Cost and pass rate are deltas against that configuration’s own baseline; the bars are how much each habit moved.
Structure-aware retrieval is the instructive failure. It does what it says — subsumed reads fall by 84.18% under Claude Code — and the bill still goes up, because querying the graph costs tokens and changes how the main agent delegates. Removing a behaviour is not the same as removing its cost.
Both skill conditions produce a markdown file the agent reads before it works. The difference is where the text comes from. Skills synthesized from traces come out trace-specific: this repository, this test runner, this file. The developer’s seven rules are trace-agnostic and read like onboarding advice — state a concrete hypothesis before you retrieve, reuse what is already in context and justify a re-read, persist and revise a script instead of creating a variant, learn the test harness, capture its output, and rerun only when the code changed or new evidence warrants it.
Now take the three rule groups apart. The paper measured the bundle, not each rule, so the split below is an attribution, not a measurement — the habit reductions are real, the cost arithmetic is ours.
Pass rates barely move in either direction: the widest swing across all twelve intervention cells is +5.33 and −2.17 points. This is a cost paper, and the right way to read it is that most of the waste was never buying accuracy.
You do not run SWE-bench. You do have a monthly figure for agent inference, and it has the same three habits inside it in roughly the same proportions, because they come from the scaffold rather than the benchmark.
Set your monthly spend and pick what you run. The percentages are the paper’s, applied to your number.
Two caveats before you spend the savings. The gains shrink on harder work: on the tougher benchmark the same developer skills land between 1.51% and 33.00%. And the effect is not uniform across models — the same file that saved 41.73% on one backbone saved 11.34% on another, so measure your own trajectories rather than the paper’s.