Coding agents open pull requests that say a thing got faster. Someone reads the diff, likes it, and merges. Nobody runs it. Out of 71,677 agent pull requests, 1,262 were performance fixes; 57% of the closed ones were merged, and when 30 merged fixes were checked out and re-run at three workload sizes, 18 delivered at least half of what they claimed, 3 fell short, 9 showed no significant gain or ran slower, and 14 changed behaviour on inputs no test covered. What predicted acceptance was not the content of the fix but the repository’s history with that agent — 31% for a newcomer against 70% with a track record.
The study starts from a public archive of agent-authored pull requests on repositories with more than a hundred stars, and narrows it down to the ones that claim a performance improvement. Two models had to agree, and then the authors had to agree with them.
Six agents wrote them, in wildly different volumes, and the fix count per agent says more about deployment than about capability. What matters for the rest of this page is that the same six agents get very different answers from maintainers.
Fifty-seven percent of closed fixes are merged, and sixty-one percent of the rejections give no stated reason at all. So what separates the two piles? The authors modelled it, and almost nothing about the change itself came out significant — not the coded content, not the description, not the tests, not the measurements included.
One diff-shape signal did survive: merged fixes delete a larger share of the lines they touch. A fix that takes code away is accepted more often than a fix that adds machinery, holding the agent and the repository fixed.
Read that as an engineer rather than as a sociologist. The signal your review process is running on is reputation, and reputation is exactly the thing that does not tell you whether this particular loop got faster.
To check a claim you need a bar, and the bar here is deliberately generous: a fix delivers if the measured improvement reaches half of what was claimed. Claim a 2× speed-up and anything at or above 1.5× counts. The authors checked out the base commit and the head commit, built both in a container with two CPUs and a thirty-minute limit, and took the median of three runs.
Set a claim and a measurement and watch the bar move. The diagonal is the line a fix has to clear.
Thirty merged fixes went through that. Eighteen cleared the bar. The interesting number is not the eighteen but the nine: fixes that a maintainer read, believed and merged, which on measurement do nothing at all or make the code slower.
The rejected pile is the control, and it is not flattering to either side: only six of the twenty-three rejected claims reproduced, with a median measured speed-up of 1.02×. Maintainers are not rejecting good fixes at any great rate. They are also not the reason the merged ones work.
Agents do write tests for these fixes — thirty-seven percent of them change at least one test file. Almost none of those tests measure the thing the pull request is about. And carrying a test, of any kind, does not raise the odds of being merged, which means the review process is not asking for one either.
The other half of the problem is the shape of the fixes. Nearly half of them are architectural rather than a local change to one hot line — a much larger share than earlier studies found for human performance fixes — and an architectural change is exactly the kind that alters behaviour on inputs the suite never exercises. Fourteen of the thirty re-run merged fixes did.
You can price this. Take the number of performance pull requests your agents open in a month and the share you merge, and apply the re-run rates to them. The rates are the paper’s; the volume is yours.
The fix is not more review. It is one required artefact: a benchmark in the pull request, a number from the base commit, a number from the head commit, produced by the same command your continuous integration can run. Eleven percent of these fixes already had it. The other eighty-nine percent were merged on a sentence.