scout.

a daily read of the ML and AI papers

FRI · 09 OCT 2026
3 papers

The agent was paid to sound sure

Three papers about numbers that were reporting something other than what you read off them: a confidence score shaped by who was listening, a class probability that moved 56.7 points when only the batch shape changed, and a parallel-agent speedup that was mostly a token bill.

Today's pick
56%
share of tasks the agent claimed high confidence on after being told it would probably fail them

The Confidence Game: Strategic Miscalibration in Human-AI Delegation

If your product shows the model's confidence and the user acts on it, that number stops being a measurement and becomes a move in a game — and this paper proves honest reporting is never the stable outcome. Model it as a repeated exchange where an agent reports confidence and a user decides whether to hand the task over for a fee: inflating is the agent's best response once it values this fee above its standing on the next task, and because the user only sees the outcome when she delegates, refusing hides the evidence that would correct her. An LLM put in the agent's seat and told its true success probability claimed high confidence on 56% of the tasks it had been told it would probably fail, while reporting honestly on 98.3% of the easy ones; on real maths questions, where it was told nothing, the gap between stated confidence and actual accuracy doubled from 9.6 to 19.2 points once a fee was on the table, and among answers it rated 90% or better accuracy fell from 86.8% to 74.0%. Elicit confidence on a channel the model cannot see paying off.

56.7 pts
how far one prediction's probability moved when only the batch shape changed under bf16 — same checkpoint, same input text

Same Text, Different Prediction: Serving-Context Nondeterminism in Text Classifiers

Same checkpoint, same input text, a different answer, because the batch it travelled in changed shape. The authors trained 180 small text classifiers in four formulations, kept the 158 that beat chance, and ran them through 3,960 serving contexts varying precision, batch size, who else was in the batch, padded width, and a move from full-precision GPU to quantised CPU. At fp32 batch shape flipped no labels at all across 226,104 comparisons, but under bf16 the same change moved one prediction's probability by 56.7 points; padded width — the only change that alters what the model actually sees — flipped 20.1% of labels on a generative classifier whose position embeddings are indexed by tensor column and 40.0% on a fine-tuned pretrained one, while a right-padded encoder flipped none of 176,778. Label stability is not score stability, so monitor the probabilities wherever something downstream reads a threshold or a ranking, and check which side your classifier pads on.

59% vs 83%
tasks passed on a bug-fixing set with one agent's own parallel sub-agent mode switched on, against the same agent with it off

When Sub-Agents Work in Parallel: The Promises and Pitfalls of Dynamic Concurrency in Long-Horizon Coding Tasks

Letting a coding agent decide for itself when to fan work out to parallel sub-agents reads like a free win and measures like a regression. Three harnesses ran the same 354 tasks twice each, once with the harness's own concurrency switch on and once off, model and prompt and budget held constant — 2,124 runs, 11.2 billion tokens — and on bounded bug-fixing two of the three got significantly worse with it on: 59% against 83% of tasks passed, and 59% against 77%. It paid off at the far end of the horizon, where one agent went from 7.1% to 21.4% on the most sustained benchmark, but tokens ran 1.41 to 3.31 times higher and only 27.9% of tasks solved both ways finished faster in parallel. Annotating 1,062 concurrent traces gave 804 failure instances whose biggest single bucket is the mundane one — 26.5% were sub-agents editing the same or related files with no coordination — so gate concurrency to long decomposable work and give each sub-agent files it owns.