Agent evaluations have grown up — they report cost and consistency now, not just a success rate. They still hold the agent itself still and vary the model. Vary the whole configuration instead, over 18,000 runs, and two things fall out. About 54% of the spread in scores is the same configuration run again, so most single-run comparisons are reading noise. And of the knobs that do matter, what the system hands the agent beats what you ask it for: pressing it in the prompt to check its answer moved real checking from 19% to 22% of runs, while giving it a checking tool roughly tripled that at every prompt level.