/p/2026-10-03 · explainer
Paper explainer · 2610.01618 · Wiedmann, Girrbach, Schmid, Akata

Half the result was the same run, twice.

Agent evaluations have grown up — they report cost and consistency now, not just a success rate. They still hold the agent itself still and vary the model. Vary the whole configuration instead, over 18,000 runs, and two things fall out. About 54% of the spread in scores is the same configuration run again, so most single-run comparisons are reading noise. And of the knobs that do matter, what the system hands the agent beats what you ask it for: pressing it in the prompt to check its answer moved real checking from 19% to 22% of runs, while giving it a checking tool roughly tripled that at every prompt level.

01 · The problem

Most of the spread is the dice, not the change

task

02 · The knobs

Five things you can change, ranked by how much they change

task

03 · The interaction

More time is only worth buying once it can be spent

how much the system tells the agent
Score against time budget, at each information level
the full protocol — the only level that clears the task

04 · The sharpest one

Asking it to check its work is not the same as letting it

how hard the prompt pushes for a check

05 · Why I care

How many runs before you believe your own A/B illustrative

how noisy your task is — borrow a profile from the paper

Results

What the paper actually measured

What it does not show

In practice