Evaluating Escalation Signals for LLM Routing: Targets, Controls, and Five Ways to Fool Yourself
You want a cheap test that says, before you pay for the big model, whether escalating this query will help at all. Semantic entropy — sample the small model ten times, cluster the answers by meaning, measure how much they disagree — is the obvious candidate, and on grade-school maths with a pair twelve times apart in size it works: it separates the small model's wrong answers from its right ones far better than a guess that reads only the question (0.871 against 0.649, where 0.5 is a coin flip and 1.0 is perfect ranking), and routing on it beats escalating at random by up to nine points of accuracy at matched cost. The same signal's strongest-looking result, on a synthetic arithmetic set, was nothing: a regex scoring question difficulty with no model in the loop matched it exactly (0.839 against 0.830), and a cheaper variant that votes on cached past outcomes fell from 0.908 to coin-flip 0.518 on a new dataset. Four cheap checks come out of this — score a question-only difficulty guess beside your signal, name which definition of "escalation worked" you mean, confirm a perfect router would gain anything at all on your data (on one benchmark the whole prize was two points), and cost the sampling honestly, because at ten samples a query this pair paid 39% more than simply always calling the large model.