Same information needed, same evidence available, same correct outcome — only the wording changes. Hint at the request instead of asking for it and a tool-using email assistant loses 22.4 points, a sandbox-checked one loses 10.7, and retrieval question answering loses 3.2. Phrase it formally and the agent stops calling tools at all on 12.2% of tasks, up from 0.4%. Two different things are going wrong underneath — padding breaks keyword retrieval before the model sees anything, while indirect and formal wording survive retrieval and show up as work the agent simply does not do — and they need different fixes.