A scripted agent did the same work with the same results, and its score dropped 0.25 — because the grader matched tool names instead of checking the operations performed. Swap in a name-aware check, rescore the same saved run, and the gap closes to zero; do the same for a genuinely sloppy agent and 0.13 of the gap stays, because that part was real. Underneath both sits the number that should govern how you read any agent result: two identical reruns already disagree about pass or fail on 14.7% to 35.6% of tasks, so a single-run improvement smaller than that is not evidence of anything.