/p/2026-09-29 · explainer
Paper explainer · 2609.31587 · Arman & Molybog

The documentation it did not need.

Two authors set out to give coding agents better context, and built the machinery properly: a benchmark that scores a description by whether code regenerated from it alone still passes the original tests, and a prompt search that found a writing style hitting full fidelity on files it had never seen. Descriptions like that are worth a great deal when the code is gone — mean share of tests passed goes from 0.08 to 0.71. Then they tested the hypothesis that started the project: better documentation helps an agent fix real repository issues. Across two model families, ten repositories and 99 tasks, no documentation condition beat the issue alone by more than one task, and the long version was the worst arm wherever it ran — while the benchmark’s own retrieved context did move the needle, so the harness could see a real effect when there was one.

01 · The measurement

Score a description by rebuilding the file from it

Most documentation quality metrics score prose against prose. This one is operational: hand a model the description and an empty project, let it write the implementation, then run the original test suite against what it produced. Tests pass, the description was faithful. That is the whole benchmark, and it turns “is this doc any good” into a number you can optimise.

What it immediately shows is that the variable everyone tunes — length — is not the variable that matters. Add facts to the description and recovery moves one for one. Make the same description seventeen times longer without adding a fact and nothing moves at all.

Five facts, two lengths
02.55

The confusable-keys experiment makes the same point adversarially. Given three near-identical setting names — max_retries, max_retry_delay, connection_max_retries — a complete verbose description matched the compact one across 28 settings, and an incomplete verbose description scored zero. Length neither helps nor hurts; the missing fact decides.

02 · The mechanism

Compress too hard and fidelity collapses

Across the eleven fixtures, description density — description tokens divided by code tokens, same tokenizer both sides — correlates with fidelity at 0.88. The shape matters more than the coefficient: every fixture that scored zero sat below 0.45 density, and every fixture at 0.9 or above cleared a third of its tests.

Move the dial to see which regime you are writing in. The paper reports the bands, not the individual fixture positions, so this widget shows the reported floor for each band rather than eleven invented points.

Density regime
0.00.51.0

So the summariser has an obligation the length budget cannot see: when the describe step squeezes a large file too hard, it drops the internal detail regeneration needs, and the description stops being a description. Completeness first, compactness second — which is exactly what the optimiser rediscovers.

03 · The method

Let the benchmark write the prompt

With a score in hand, the describe prompt becomes a search problem. A proposer model suggests new describe prompts, each is scored on the fixtures, and the objective mildly penalises length: J = mean fidelity − 0.1 × (mean words / 300). The discovered prompt asks for precise imports, module-level constants with their literal values, exact signatures and defaults, and every exception and return condition — independently rediscovering each failure mode the authors had found by hand.

The objective, as the search sees it

Run that loop and the numbers are genuinely good: full fidelity on the three training fixtures, and on files it never optimised against, held-out fidelity rises from 0.5 to 1.0.

What the discovered prompt bought

At this point the project has a working doc pipeline: measurable, optimised, generalising. Everything above is the setup for the result the authors actually went looking for.

04 · The result

Now let the agent read the file

The downstream test is the ordinary one: an agent, an issue, a repository, does it resolve the task. Four conditions — the issue alone, the issue plus a compact description of about eighty words, the issue plus the full-length optimised description, and the issue plus the benchmark’s own retrieved context from past tasks.

Flip the switch. The only thing that changes is whether the agent can open the source.

The same descriptions, two worlds

Read the control row before you dismiss the harness: on the same repository and the same 29 paired tasks, the benchmark’s own retrieved context resolved 15 against 14 for the issue alone, reproducing the effect its authors report. The evaluation can see an improvement. Documentation does not produce one — and the full-length description, the very artefact the optimiser worked so hard for, is the worst arm wherever it ran.

The authors’ explanation is the sentence worth keeping: a description faithful enough to rebuild a file is a restatement of that file. Inside the context window it is redundant, and occasionally harmful, because a complete account of a readable file nudges the agent to rewrite rather than patch.

05 · Your own doc layer

Where the compression is actually worth paying for

The useful reframe: documentation is a compression format for code that will not fit. That tells you exactly where to spend. A repository your agent can page through does not need a summary layer. A partner’s API you only see through a client library, a binary, a 400-file service the agent gets 30 files of — those are the compression regimes.

Set how much of the relevant code your agent can actually read, and what your doc layer costs to keep current.

Should you build the doc layer? illustrative
expected uplift
what you are paying for it

And if you do build it, score it the way this paper does. Regenerate the thing from your description and run the tests. A doc that cannot rebuild what it describes will not help an agent that cannot see the original either.

Results

What the paper actually measured

What it does not show

In practice