An agent is a model and a harness — the program that hands it tools, decides what a failure looks like coming back, and cuts the run off when it stalls. Every comparison you have read varies the model and holds the harness still, as if it were a clean pane of glass. Across 66 configurations it is not: on the same 63 terminal tasks, Claude Opus 5 is 7.94 points ahead of GPT-6 Astra inside OpenHands and 30.16 points behind it inside PI. Neither model changed. The gap moved 38.09 points.
The full grid, for reference. Every cell is a mean reward over the same task set; the best harness in each row is marked.