scout.

a daily read of the ML and AI papers

THU · 08 OCT 2026
3 papers

The router was reading the question

Three papers about the layer you wrote rather than the model you called: an escalation signal that turned out to be reading the question, agents that judge an action correctly and then take it before the evidence is in, and a transcript that halves when the harness stops re-sending itself.

Today's pick
+9 pts
more correct answers than escalating at random for the same spend — on the one benchmark where the signal survived every check

Evaluating Escalation Signals for LLM Routing: Targets, Controls, and Five Ways to Fool Yourself

You want a cheap test that says, before you pay for the big model, whether escalating this query will help at all. Semantic entropy — sample the small model ten times, cluster the answers by meaning, measure how much they disagree — is the obvious candidate, and on grade-school maths with a pair twelve times apart in size it works: it separates the small model's wrong answers from its right ones far better than a guess that reads only the question (0.871 against 0.649, where 0.5 is a coin flip and 1.0 is perfect ranking), and routing on it beats escalating at random by up to nine points of accuracy at matched cost. The same signal's strongest-looking result, on a synthetic arithmetic set, was nothing: a regex scoring question difficulty with no model in the loop matched it exactly (0.839 against 0.830), and a cheaper variant that votes on cached past outcomes fell from 0.908 to coin-flip 0.518 on a new dataset. Four cheap checks come out of this — score a question-only difficulty guess beside your signal, name which definition of "escalation worked" you mean, confirm a perfect router would gain anything at all on your data (on one benchmark the whole prize was two points), and cost the sampling honestly, because at ten samples a query this pair paid 39% more than simply always calling the large model.

95–99% vs 52%
how often agents judged the same action correctly on paper, against how often they actually carried it out with the required evidence established first

From Evidence to Action: How Tool-Using Agents Fail

Your agent correctly says “yes, cancelling this order is the right call” when you show it the situation on paper, then goes and cancels an order it never looked up. A team built 656 cases across customer operations, infrastructure, legal and financial work, research, smart-home control and healthcare, ran them under ten model-and-harness combinations, and scored each episode with a ledger recording which facts were established and when — so acting before you looked is caught even when the outcome happens to be right. On the same actions, judging on paper scores 95–99% while executing with the evidence in first peaks at 52%: agents either stopped before finishing the investigation (21.7–62.9% of cases) or fired the action before the required reads landed (37.0–66.9%). Once the evidence was in, single actions almost always succeeded (83.3–100%) — so the thing to instrument is not the action, it is whether the reads it depends on happened first.

2.5×
fewer tokens per session at best across five long-horizon benchmarks, with task resolution unchanged

ReFold: Training-Free Reversible Inter-Turn Context Folding for Long-Horizon Agents

A long-running agent re-sends its whole transcript at every step, so the file it printed at step 4 is still being paid for at step 40. This is a rendering layer in front of an ordinary reason-and-act loop that compresses only what the model sees and leaves the real history untouched: content an earlier turn already displayed becomes a one-line stub pointing at the turn that still holds it, and turns the agent itself marks finished fold into a single progress note. Nothing is predicted and nothing is lost — a wrong fold costs one restore command — and the rendered prefix is rewritten every three steps rather than every step, keeping the prefix-cache hit rate at 87% against the 51% of rewriting continuously. Across five long-horizon benchmarks it cut tokens by 18–53% on a dense model and up to 60% on a mixture-of-experts one, with task resolution unchanged, at 0.4 ms of overhead a step against 3.3 s for summarising the transcript instead.