A classifier is supposed to be a function of its input. In practice it is a function of its input and of the batch it travelled in, the width it was padded to, the precision it was served at, and the device it landed on. 180 small classifiers were trained and the 158 that beat chance pushed through 3,960 serving contexts. At fp32 a change of batch shape flipped no labels at all across 226,104 comparisons; under bf16 the same change moved one prediction’s probability by 56.7 points. The largest effect is the dullest: padded width, the one change that alters what the model actually reads, flipped 20.1% of labels on a generative classifier with learned absolute positions and 40.0% on a fine-tuned pretrained one, while a right-padded encoder flipped none of 176,778. The sentence to keep: label stability does not imply score stability, and every threshold downstream reads the score.