/p/2026-10-04 · explainer
Paper explainer · 2610.00054 · Villuri, Shaik, Doboli

The first token is not the verdict.

Reading a judge’s verdict out of the logits of its first token costs nothing, needs no generation, and is exactly what a likelihood-scoring harness hands you. It also answers a different question. Judges often do not lead with a verdict token at all, and a forced read on those pairs returns whichever response was shown first: pooled over the 924 pairs where the judge did not commit, the forced read changes its answer when the two responses are swapped 89.7% of the time, against 47.5% when the judge is allowed to finish the sentence. What moves is the audit, not the verdict.

01 · The problem

Two readouts, two different position-bias numbers

02 · The mechanism

The judge was not answering yet

judge, dataset and verdict letters

the two characters you ask for

03 · What it distorts

It moves the audit, not the verdict

highlight one condition
Flip-rate gap against how often the judge leads with a verdict
the forced read runs ahead of the generated one the two readouts agree

04 · The second failure

It opens with one letter and argues for the other

05 · Why I care

How much of your position-bias figure is the readout illustrative

Results

What the paper actually measured

What it does not show

In practice