/p/2026-09-30 · explainer
Paper explainer · 2609.32035 · Althoubi

Wrong answers that agree.

Self-consistency rests on a claim about disagreement: sample the model several times, and if it is unsure the wrong answers will scatter, so agreement is evidence of being right. Holding the weights fixed and flipping only a thinking-mode flag, over five benchmarks and 74,944 samples, the chance that two independently drawn wrong answers match went up in all ten dataset-and-size comparisons. On open-ended maths, reasoning cut the number of distinct answers the model produces to between 0.43 and 0.65 of the non-reasoning count. Voting still works, for a reason that should worry you — reasoning also shrinks the set of problems where extra samples could decide anything, by 2.7×. What does not survive is confidence: of 280 weighted-voting combinations, none beat a plain majority.

01 · The assumption

Agreement is only evidence if disagreement was possible

Sample a model eight times, keep the answer that came up most often, ship that. The reason this works is supposed to be that a model which does not know scatters: there are many wrong answers and only one right one, so wrong samples spread out and cannot out-vote the correct ones even when the correct ones are a minority. The whole method leans on wrong answers being uncorrelated.

Drag the concentration of the wrong answers and watch the vote change. The model is equally wrong in every position on the slider — five of its eight samples are wrong either way. Only where those five land moves.

Eight samples, one vote illustrative

The collision index on the slider is the paper’s own measure: the probability that two independently drawn wrong answers to the same problem are the same answer. Nothing about the model’s accuracy changes as you move it. What changes is whether a minority of correct samples can still win.

02 · The measurement

Same weights, one flag, errors pull together

The trick that makes this clean: one model family serves thinking and non-thinking answers from identical weights, chosen by a flag in the chat template. So nothing about the model differs between the two arms — not the parameters, not the tokenizer, not the fine-tuning. Only whether a chain of reasoning gets generated before the answer.

Every model answered every problem eight times at temperature 0.6. Pick a cell to see what the flag did to it.

Collision among wrong answers, per benchmark and size
benchmark & model size

All ten cells

Ten of ten in the same direction, which is a one-in-a-thousand coincidence if the flag did nothing. And the effect survives the obvious objection — that the reasoning arm is simply more accurate, so its errors are the hard problems where everyone agrees on the same trap. Restricted to only the problems each arm itself gets wrong, nine of nine cells still favour reasoning.

03 · The mechanism

Reasoning shrinks the answer space it is sampling from

Where the answer can be any number, you can count the distinct answers a model produces across a whole benchmark. Reasoning produces far fewer of them — 339 where the same weights without reasoning produced 522 — and the ratio sits between 0.43 and 0.65 everywhere it can be measured. A chain of thought is a commitment device: once the model has written three paragraphs towards an answer, the next sample from the same prompt tends to write the same three paragraphs.

Where the answer space is bounded — four options, ten options — both arms have exactly the same option set available, and reasoning piles mass onto one wrong option instead. No positional or format bias explains that at fixed weights.

Distinct answers produced, open-ended benchmarks
benchmark & model size
0275550 distinct answers

This is the part to carry into your own system. If you sample a reasoning model several times for diversity — candidate plans, candidate SQL, candidate extractions — you are getting less diversity than the temperature suggests, and the duplicates are not evidence.

04 · Why the bill is small anyway

The vote survives because there is less left to vote on

If concentrated errors were the whole story, majority voting over a reasoning model would be visibly worse. It is not, and the reason is worth understanding rather than being relieved about. Voting can only change an answer on problems where fewer than half the samples are correct — three or fewer of eight. Call that the decidable set. Reasoning shrinks it too.

Pick a cell to see how much smaller.

Problems where extra samples could decide anything
benchmark & model size
0%20%40% of problems

So the two effects cancel: errors agree more, and fewer problems are left where that matters. The accuracy gain from eight samples instead of one looks three times larger for the non-reasoning arm, but that arm also has three times as much room — expressed as the fraction of remaining headroom each converts, it is 0.250 against 0.249. Identical. The mechanism is real and its aggregate cost hides.

05 · What breaks instead

The confidence number you were going to weight by

If agreement is a weaker signal than it looks, the natural move is to weight the vote — by log-probability, by self-reported confidence, by output length, by a learned combination. Seven such methods, eight models, five benchmarks: 280 combinations, and after correcting for running that many tests, not one beats a plain majority. Five of the 280 clear p<0.05 uncorrected, against fourteen you would expect by chance alone.

The reason is mundane. Weighted voting almost never disagrees with the majority, and when it does it is barely better than a coin.

What weighting your vote buys illustrative
decisions the weighting changes
net answers gained

And one signal does something worse than being useless. Flip the reasoning flag and the fitted weight on the answer token’s log-probability changes sign: with thinking off, higher confidence at the answer token predicts a correct answer; with thinking on, it predicts a wrong one. Once a chain has committed the model to an answer before the answer token is emitted, that log-probability reports how firmly the chain committed — and among the problems it got wrong, firmer is worse.

Fitted weight on answer log-probability

A learned six-signal combination makes the point cleanly: pooled over every dataset including the one it was fitted on it gains +0.0118; restricted to the four it was not fitted on it gains −0.0003, losing on five of eight models. A signal good enough to rank is not automatically good enough to route.

Results

What the paper actually measured

What it does not show

In practice