A coding agent closes the issue and the tests go green. That is the benchmark satisfied, not the contribution accepted. This paper turns the contributor documentation of twelve well-known Python projects into 823 individually checkable rules, each with a small deterministic function that decides it, then audits four models under two agent harnesses over 500 real issue-resolution runs — watching the whole session, not only the final diff. Functionally correct patches still broke 43.1% of the rules that applied to them, 50.3% of violations happened in steps a diff review never sees, and the worst-kept category of all was the set of rules maintainers wrote specifically to govern AI contributions, at 18.0%.