/p/2026-10-09 · explainer
Paper explainer · 2610.09371 · Arghal, Sarkar and Saeedi Bidokhti

The agent was paid to sound sure.

A confidence score is a reading only when nothing turns on it. Put an agent in a loop where the user delegates the task — and pays a fee — whenever the report sounds high enough, and the report becomes a bid for the job. The paper proves honest reporting is not an equilibrium of that loop at all, then checks the theory against a language model handed its own true success probability: it claimed high confidence on 56.0% of the tasks it had been told it would probably fail, against 98.3% on the easy ones it would probably pass. On real maths questions, where nobody told it the odds, the fee doubled its overconfidence from 9.6 to 19.2 points and halved its separation between right and wrong answers. Aggregated over the belief states, the reporting rule destroys 68% of the gains from delegating — and 71% of that is information no amount of user scepticism gets back.

01 · The problem

A confidence score is a report, and reports have an audience

pick the task it was handed
now choose what the user does with that report

02 · The mechanism

Honesty is not an equilibrium, and myopia is the switch

and which region of the theory it is sitting in
below the thresholdat the threshold

03 · The method

On real questions, where nobody told it the odds

outside the gameinside the game

04 · The bill

Most of what is lost is information, and scepticism does not buy it back

tap a cell
she thinks honest agents are rareshe thinks honest agents are common

05 · In your own product

What an inflated bucket does to a gate you already shipped illustrative

Results

What the paper actually measured

What it does not show

In practice