Finding the Right Fit: Model-Harness Interactions across Agent Tasks
Choosing an agent stack means choosing two things — the model, and the harness that hands it tools, errors and context — and the second is usually treated as plumbing. Across 66 model-and-harness pairings the rankings do not survive the swap: on 63 terminal tasks one frontier model leads another by 7.94 points inside OpenHands and trails it by 30.16 inside the leaner PI, and for four of the five models the best harness changes from one benchmark to the next. What separates the harnesses is mundane — whether a failure comes back as something the model can act on: one has no shell timeout and left 34 runs silently blocked in a command that never returned, where another capped commands at 300 seconds and turned the same hang into a readable error. A model bake-off run inside someone else's harness is telling you about their harness, so run yours inside yours.