/p/2026-09-30 · explainer
Paper explainer · 2609.35357 · Ma, Wang, Huang, Wang, Wang, Si, Wei & Zhang

The agent stopped looking.

Give a coding agent a feature in five instalments on a real repository, with the workspace carrying over, and it explores less each time. On the first turn it reads 83.6% of the existing functions it was meant to call; by the fifth, 35.4%. It also stops reusing what it wrote itself, 83.9% down to 69.1%, although that code is still sitting in front of it. What accumulates is duplicated logic in 50.8% of chains by turn five against 13.8% at turn one — and across 3,000 turns, the tests never notice, holding near 91% throughout.

01 · The question

A passing test cannot see a second implementation

Every benchmark a coding agent is measured on asks the same thing: did the tests go green. That question is blind to a whole category of damage. If the agent writes its own date parser next to the one three modules over, the tests pass, the diff looks fine, and you now have two date parsers — so the next bug in that logic is a fix applied twice, and nothing in your pipeline flagged it.

The benchmark here reveals a requirement one turn at a time, keeps the workspace between turns, and measures three things beyond the pass rate: how much of the relevant existing code the agent read, how much of it the agent called, and how many chains end up carrying a re-implementation of something they already had.

First turn against fifth

What moves between turn 1 and turn 5

Across the eight model-and-harness pairs the declines are not a quirk of one setup: reading of relevant repository code falls by 47.5% to 61.6%, reuse of the agent’s own earlier code by 8.1% to 24.0%, and the share of chains carrying duplicated logic rises by 22.6 to 48.0 points.

02 · Two different failures

It never looked, or it looked and wrote its own anyway

Separating reading from calling is what makes this diagnosable. If the agent never read the function, no amount of prompting about reuse will help — the problem is exploration. If it read the function and wrote its own regardless, exploration is fine and the problem is the decision. The two need different fixes, and this benchmark tells them apart by measuring recall and reuse on the same targets.

Switch between the two failures to see which applies where.

Where the reuse is lost
the gap
0%50%100%

Both happen, in different places. On the repository’s own code the failure is exploration: by turn five the agent has read barely a third of what it needed, so most of the loss is code it never saw. On its own earlier code the failure is the decision: recall there is near saturated on every turn — the code is in the workspace and the agent knows it — and reuse still falls to 69.1%.

03 · What accumulates

Redundancy compounds while the dashboard stays green

Duplication is not a per-turn error rate, it is a stock. Once a second implementation exists it stays, and the next turn writes on top of it. By the fifth turn half the chains carry one — and the resolve and pass rates over the same runs sit inside a band a couple of points wide.

Move the turn. Turns 1 and 5 are measured; the three between them are a straight line drawn between those endpoints, marked as such.

Accumulated duplication by turn illustrative between the endpoints

0%50%100%

That last bar is the argument for the whole paper. The strongest configuration resolved between 66.4% and 72.3% of tasks and passed 90.8% to 91.6% of tests, and it was accumulating duplicates the whole way. If pass rate is your only gate, this failure mode is invisible by construction.

04 · The fix that worked

An index of what exists, not a transcript of what happened

One ablation is the practical payload. Holding the model and harness fixed, the agent was given three different kinds of memory of its own earlier work: none, the full source of what it had written, and an interface-level index — the signatures and what they do, without the bodies. Only one of them changed anything.

Pick a memory setting.

Reuse of the agent’s own earlier code
what the agent is given about its own past work
0%50%100% reused

Source-level memory is the version most people build, and it did nothing: 29.2% against 30.0% with no memory at all. Handing a model a pile of its own code is not the same as telling it what it already has. The index more than doubled reuse, to 67.8%, and resolve rates across all three settings stayed inside 27.7% to 30.4% — the memory changed the redundancy and not the correctness, which is exactly the trade you want and exactly the one a pass-rate gate cannot see.

05 · Your own repository

How many fixes you will be applying twice

You do not run this benchmark, but you do run agents over several turns on one codebase, and the shape carries: the first instalment explores, the later ones stop. The numbers below apply the measured duplication rates to your own volume of agent-authored work.

Set how many multi-turn chains your team runs a month and how long they get.

A month of multi-turn agent work illustrative
chains carrying a duplicate
with an interface index in place

The second meter deserves a caveat the paper is explicit about: the interface-index result is measured on reuse of the agent’s own earlier code, not on the repository’s. Applying it to the duplication rate is an extrapolation, which is why it is labelled illustrative. The part that is not an extrapolation is that nothing in your test suite will tell you either number.

Results

What the paper actually measured

What it does not show

In practice