/p/2026-10-09 · explainer
Paper explainer · 2610.09111 · Kasa, Kasa and Negi

Same text, a different prediction.

A classifier is supposed to be a function of its input. In practice it is a function of its input and of the batch it travelled in, the width it was padded to, the precision it was served at, and the device it landed on. 180 small classifiers were trained and the 158 that beat chance pushed through 3,960 serving contexts. At fp32 a change of batch shape flipped no labels at all across 226,104 comparisons; under bf16 the same change moved one prediction’s probability by 56.7 points. The largest effect is the dullest: padded width, the one change that alters what the model actually reads, flipped 20.1% of labels on a generative classifier with learned absolute positions and 40.0% on a fine-tuned pretrained one, while a right-padded encoder flipped none of 176,778. The sentence to keep: label stability does not imply score stability, and every threshold downstream reads the score.

01 · The problem

Four things you change without thinking of them as inputs

pick a serving change

02 · The mechanism

The label held and the score moved 56.7 points

now set the gap between the top two classes

03 · The method

Padding is not arithmetic — it changes what the model reads

pick a model family
satisfies the shift conditionmeasured flip rate

04 · The fix

One mitigation per mechanism, and what each one costs

tap a mitigation

05 · In your own service

How many decisions a day the near-tie tail is rewriting illustrative

Results

What the paper actually measured

What it does not show

In practice