Before paying for the big model you want a cheap signal that says whether escalating this query will help. Semantic entropy is the obvious candidate, and on grade-school maths with a pair twelve times apart in size it earns its place: a ranking score of 87.1 against 64.9 for a guess that never calls a model, and nine points of accuracy over escalating at random for the same spend. On a synthetic arithmetic ladder the same signal’s best-looking number was nothing at all — a regex reading only the question matched it exactly, 83.9 against 83.0 — and at ten samples a query the router cost 39% more than simply always calling the large model. The paper’s real output is five checks that catch all of this before it gets reported as real.