/p/2026-10-06 · explainer
Paper explainer · 2610.02627 · Chen, Dutt, Taheri and Williams

Your eval asks each question once.

Same information needed, same evidence available, same correct outcome — only the wording changes. Hint at the request instead of asking for it and a tool-using email assistant loses 22.4 points, a sandbox-checked one loses 10.7, and retrieval question answering loses 3.2. Phrase it formally and the agent stops calling tools at all on 12.2% of tasks, up from 0.4%. Two different things are going wrong underneath — padding breaks keyword retrieval before the model sees anything, while indirect and formal wording survive retrieval and show up as work the agent simply does not do — and they need different fixes.

01 · The problem

Ten ways to ask the same thing, and one of them is in your eval

how the request was rewritten

02 · Two different failures

Lost before the model, or lost after it had everything it needed

highlight one case

03 · In the agentic setting

It does not do the wrong thing. It does nothing.

which system, rewritten which way

04 · Does it hold up

Three models, two benchmarks, the same one condition that always hurts

pick a rewrite to read across

05 · In your own product

Build the variant suite and see what it costs you illustrative

tap a rewrite to add it to your suite

Results

What the paper actually measured

What it does not show

In practice