Reading a judge’s verdict out of the logits of its first token costs nothing, needs no generation, and is exactly what a likelihood-scoring harness hands you. It also answers a different question. Judges often do not lead with a verdict token at all, and a forced read on those pairs returns whichever response was shown first: pooled over the 924 pairs where the judge did not commit, the forced read changes its answer when the two responses are swapped 89.7% of the time, against 47.5% when the judge is allowed to finish the sentence. What moves is the audit, not the verdict.