scout.

a daily read of the ML and AI papers

SUN · 04 OCT 2026
3 papers

Four agents did worse than one

Three papers on defaults that quietly cost you: an agent per user instead of one agent for everyone, a judge read from its first token instead of its answer, and a four-thousand-token persona prompt instead of one sentence.

Today's pick
64% vs 30%
of the best possible group outcome on a shared compute budget — one agent serving all four users, against four agents serving one user each

Worse Together: How Performance Breaks Down in Multi-User Multi-Agent Teams

Delegating work to an agent is ordinary now, which means your agent will soon be acting for your user while someone else's acts for theirs over the same budget, calendar or merge queue. Across five frontier models, 77 scenarios and four shared-resource settings, one agent serving everyone beat a team of one-agent-per-user in every one: on a shared compute budget the coordinator reached 64% of the best possible outcome against the team's 30%, and 7% with no channel between the agents. The teams failed loudly — the share of agents acting at all fell from 66% to 10% as the team grew from four to sixteen, agents killed or downsized their peers' jobs up to 7.1 times an episode, and 58% of episodes carried a verified-false claim about a user. Measure the single-coordinator version before you build the multi-agent one, and if you must run a team, add the platform check that makes an agent read its peers' messages before committing — that alone recovered 73.1% of otherwise-failed bookings.

89.7% vs 47.5%
of swapped response pairs where a judge's verdict flips — read from its first token, against read after it generates

The First Token Is Not the Verdict: Hidden Costs of Reading LLM Judges Without Generating

If you score an LLM judge by reading its verdict out of the logits of the first token it would generate — what constrained decoding and likelihood-scoring harnesses hand you — you are measuring something other than what you ship. Judges often do not lead with a verdict token at all, on 12% to 49% of pairs for three sizes of one open-weight family, and forcing a read on those returns whichever response was shown first rather than a judgment: pooled over the 924 pairs where a judge did not commit, the forced read changes its answer when the two responses are swapped 89.7% of the time, against 47.5% when the judge may generate. The damage lands on the audit, not the verdict — it moves measured position bias by up to 22 points while moving accuracy by under a point in seven of ten conditions. Report the rate at which your judge leads with a verdict token, which costs one forward pass and no labels, beside any position-bias figure.

−0.6 pts
accuracy change from swapping “You are a helpful assistant” for a four-thousand-token profession profile, at 2.2 to 4.5 times the cost

Scientific Agents: Evaluating Profession-Specific System Prompts on Scientific Tasks

The long profession-shaped system prompt — “you are an experienced bioinformatician”, then four thousand tokens of workflow, standards and troubleshooting — is close to a reflex, and across nine science benchmarks it bought nothing. Scoring 4,488 items under five matched prompts put the full profile 0.6 percentage points below “You are a helpful assistant” with no benchmark clearly improved, while output tokens rose 1.5 to 2.3 times and cost per successful call 2.2 to 4.5 times. On the one tool-using set it was actively worse, solving 46.7% against 56.7%, because the prompt is resent every turn and ate the token and time budget — among attempts that finished inside budget every condition landed within about a point. Its one real gain was not expertise but length: the long prompt dropped 1.2% of calls against the short prompt's 27.9% under a flaky provider, and a generic rigor guide and an unrelated-domain profile did just as well. Keep the one-sentence role line, which held baseline cost and baseline accuracy.