scout.

a daily read of the ML and AI papers

WED · 30 SEP 2026
3 papers

It did what the error message said

Three papers on instructions that quietly mislead: the error text your tool server hands back, the agreement you are reading as confidence, and the repository code your agent stopped looking at.

Today's pick
45% → 84%
of turns recovered after an expired-credential error, once the step names the server's own login tool instead of a terminal command

MCP Error Messages Written for Developers Hurt the Most Capable Agents Most

Your tool server wraps a web API written for human developers, so when a token expires it returns something like "run: reddit-mcp-buddy --auth" — and the agent, which can only call tools, relays that to you and stops. Across 150 widely used servers, 949 of 3,001 error messages tell the caller what to do next and about half of those steps depend on something the server cannot see: on expired credentials 62 of 67 ask for a terminal command, a config edit or a web page. In 15,120 trials on five tool-only models the terminal-command step left 45% of turns recovered, and the damage scaled with capability — 18 points of lost recovery on the oldest model against 69 on the largest newest one, which made 0.60 tool calls per trial and handed the repair back to the user with the login tool sitting in its list. Both fixes are one line: name the server's own tool in the step (84% on credentials, 88% on rate limits once the message names the call to repeat), or delete the instruction sentence with a prompt before the model reads it (82%, nine cents across all 949 messages).

0 of 280
confidence-weighted voting schemes that beat plain majority voting, across eight models and five benchmarks

Reasoning Concentrates Errors, and Self-Consistency Never Notices

Sampling a model several times and taking the majority answer rests on the idea that an unsure model scatters its wrong answers, so agreement means correctness. Holding weights fixed and flipping only a model's thinking mode over five benchmarks and 74,944 samples, this paper shows reasoning makes wrong answers agree with each other: the chance that two independently drawn wrong answers match rose in all ten dataset-and-size comparisons, and on open-ended maths reasoning cut the distinct answers produced to 0.43–0.65 of the non-reasoning count. The aggregate cost is milder than that sounds, because reasoning also shrinks the set of problems extra samples could decide by 2.7 times — but weighting by confidence recovers nothing: of 280 method-dataset-model combinations none beat plain majority voting, and the log-probability of the answer token flipped sign, predicting correctness with thinking off and error with it on. Score a confidence signal by the decisions it changes rather than by how well it ranks, and re-check any signal you tuned before you turned reasoning on.

50.8%
of five-turn task chains end carrying duplicated logic, against 13.8% after the first turn, while pass rates stay flat

Do Coding Agents Reuse Existing Code or Reinvent the Wheel?

Hand a coding agent a feature request in five instalments on a real repository and it stops looking around: it reads 83.6% of the existing functions it was meant to call on the first turn and 35.4% by the fifth, and reuses its own earlier code less too, 83.9% down to 69.1%, though that code is still in the workspace. What accumulates is duplicated logic in 50.8% of chains by turn five against 13.8% at turn one, while pass rates sit near 91% — a second implementation of something you already have breaks no test, so nothing you gate on will tell you. Across 3,000 turns on five Python libraries with four models under two harnesses the pattern held throughout, and one ablation is worth copying: giving the agent an interface-level index of what it had already written lifted reuse of its own code from 30.0% to 67.8% with correctness untouched. Feed multi-turn agents a symbol index rather than raw history, and review structurally against the call graph, because your test suite is blind to exactly this.