/p/2026-10-07 · explainer
Paper explainer · 2610.06193 · Truong, Goh, Le-Cong and Huo

Green tests, rejected patch.

A coding agent closes the issue and the tests go green. That is the benchmark satisfied, not the contribution accepted. This paper turns the contributor documentation of twelve well-known Python projects into 823 individually checkable rules, each with a small deterministic function that decides it, then audits four models under two agent harnesses over 500 real issue-resolution runs — watching the whole session, not only the final diff. Functionally correct patches still broke 43.1% of the rules that applied to them, 50.3% of violations happened in steps a diff review never sees, and the worst-kept category of all was the set of rules maintainers wrote specifically to govern AI contributions, at 18.0%.

01 · The problem

Passing the tests and being mergeable are not the same measurement

tap a model and harness

the twelve projects, by tasks drawn from each
02 · The mechanism

Half the violations never reach the diff

03 · The method

Eight hundred and twenty-three rules, one function each

tap a category of rule

04 · The obvious fix, tested

Handing them the rulebook helps, and does not solve it

does the agent even end up holding the rules

05 · In your own repo

What reaches review in a week of agent pull requests illustrative

Results

What the paper actually measured

What it does not show

In practice