/p/2026-10-01 · explainer
Paper explainer · 2609.37985 · Qi, Li, Chen et al.

A merge is not a measurement.

Coding agents open pull requests that say a thing got faster. Someone reads the diff, likes it, and merges. Nobody runs it. Out of 71,677 agent pull requests, 1,262 were performance fixes; 57% of the closed ones were merged, and when 30 merged fixes were checked out and re-run at three workload sizes, 18 delivered at least half of what they claimed, 3 fell short, 9 showed no significant gain or ran slower, and 14 changed behaviour on inputs no test covered. What predicted acceptance was not the content of the fix but the repository’s history with that agent — 31% for a newcomer against 70% with a track record.

01 · The set

Every agent pull request that says something got faster

The study starts from a public archive of agent-authored pull requests on repositories with more than a hundred stars, and narrows it down to the ones that claim a performance improvement. Two models had to agree, and then the authors had to agree with them.

From every agent pull request to the study set
filtering stage
0

Six agents wrote them, in wildly different volumes, and the fix count per agent says more about deployment than about capability. What matters for the rest of this page is that the same six agents get very different answers from maintainers.

Performance fixes per agent
0250500 fixes

02 · What gets merged

Acceptance tracks who is asking, not what they changed

Fifty-seven percent of closed fixes are merged, and sixty-one percent of the rejections give no stated reason at all. So what separates the two piles? The authors modelled it, and almost nothing about the change itself came out significant — not the coded content, not the description, not the tests, not the measurements included.

What actually moves acceptance

One diff-shape signal did survive: merged fixes delete a larger share of the lines they touch. A fix that takes code away is accepted more often than a fix that adds machinery, holding the agent and the repository fixed.

Merge rate by agent
agent
0%50%100% merged

Read that as an engineer rather than as a sociologist. The signal your review process is running on is reputation, and reputation is exactly the thing that does not tell you whether this particular loop got faster.

03 · The re-run

Half the claim, or it did not deliver

To check a claim you need a bar, and the bar here is deliberately generous: a fix delivers if the measured improvement reaches half of what was claimed. Claim a 2× speed-up and anything at or above 1.5× counts. The authors checked out the base commit and the head commit, built both in a container with two CPUs and a thirty-minute limit, and took the median of three runs.

Set a claim and a measurement and watch the bar move. The diagonal is the line a fix has to clear.

The delivery criterion, live

Thirty merged fixes went through that. Eighteen cleared the bar. The interesting number is not the eighteen but the nine: fixes that a maintainer read, believed and merged, which on measurement do nothing at all or make the code slower.

Re-running what was merged, and what was rejected
0

The rejected pile is the control, and it is not flattering to either side: only six of the twenty-three rejected claims reproduced, with a median measured speed-up of 1.02×. Maintainers are not rejecting good fixes at any great rate. They are also not the reason the merged ones work.

04 · Why nothing catches it

The tests are there, and they are not measuring speed

Agents do write tests for these fixes — thirty-seven percent of them change at least one test file. Almost none of those tests measure the thing the pull request is about. And carrying a test, of any kind, does not raise the odds of being merged, which means the review process is not asking for one either.

What the fixes carry, and what it covers
0%50%100%

The other half of the problem is the shape of the fixes. Nearly half of them are architectural rather than a local change to one hot line — a much larger share than earlier studies found for human performance fixes — and an architectural change is exactly the kind that alters behaviour on inputs the suite never exercises. Fourteen of the thirty re-run merged fixes did.

What the slow code was doing, and how deep the fix went

05 · Your own repo

What a month of agent performance work is actually worth

You can price this. Take the number of performance pull requests your agents open in a month and the share you merge, and apply the re-run rates to them. The rates are the paper’s; the volume is yours.

A month of merged speed-ups illustrative
merged and working as claimed
merged on a claim that does not hold

The fix is not more review. It is one required artefact: a benchmark in the pull request, a number from the base commit, a number from the head commit, produced by the same command your continuous integration can run. Eleven percent of these fixes already had it. The other eighty-nine percent were merged on a sentence.

Results

What the paper actually measured

What it does not show

In practice