Tool servers wrap web APIs written for human developers, so their errors tell the reader to run a command, edit a config or open a page. An agent that can only call tools reads the same sentence and obeys it — by telling you to go and do it. In 150 widely used servers, 949 of 3,001 error messages give the caller a next step, and on expired credentials 62 of 67 of those steps point somewhere the agent cannot go. That one sentence left 45% of turns recovered, and it cost the newest, largest model 69 points where it cost the oldest 18. Name the server’s own login tool instead and recovery is 84%; delete the sentence before the model reads it and it is 82%.
An error message is a piece of interface design, and almost every tool server inherited its design from the web API underneath. That API was written for a person at a keyboard: run this, set that, open this page, wait a bit. The agent calling your server has a tool list and nothing else — no shell, no browser, no config file it can edit, and no way to wait out a window and come back.
The survey read every distinct error string in 150 servers (426 to 186,635 stars, median 1,609) and asked two questions of each: does it tell the caller what to do next, and does that step depend on something the server cannot see about the caller?
Read the credential row twice. Sixty-two of the sixty-seven servers that bother to help you on an expired token help in a way an agent cannot use, and the five that name a tool are the ones whose messages were written after somebody watched an agent read them.
To separate the text from everything else, the experiment fixes the task and varies only the error string. It replays a multi-turn tool-use task up to the call that matters, swaps that call for a failing one, hands back one of six versions of the same error, and then lets the agent run until the turn ends or it has made eight tool calls. Recovered means the world ended up where the reference solution would have left it — every service’s state matches.
Step through the six texts. Everything except the sentence in the message is identical.
The ordering is the finding. Saying nothing useful (61%) beats pointing at a terminal (45%). Naming the wrong tool (81%) also beats pointing at a terminal, because a wrong tool call still leaves the agent inside the loop where its next move can be right. The instruction is not weak; it is strong and it points out of the room.
Five models from three generations read the same message. The generic notice and the bare cause land within a few points of each other across all of them. The terminal command does not: the loss it causes grows monotonically with how well the model follows written instructions, from 18 points on the oldest to 69 on the largest newest one.
Pick a model, then flip the message between the original step and the same message with the step removed.
The largest model is not confused. In 48 of its 68 unrepaired trials it wrote a clear final message asking the user to reconnect or sign in, or pointing at the command — with the login tool listed in front of it the whole time. It treated the sentence as a specification of who does the work, which is exactly what a well-written developer-facing error is.
Both remedies are edits to a single sentence, and they are available to different people. If you ship the server, replace the human step with the tool the agent has: “Call ticket_login first” instead of the command, and “Wait a few seconds and call get_issue again” instead of a bare wait. If you ship the agent, strip the step before the model sees it and keep only what went wrong.
Choose whose code you control, and which failure you are looking at.
The rate-limit case is the largest single swing in the paper: GitHub’s “Wait before retrying.” leaves 6% recovered because the agent waits and then has nothing to do; naming the call to repeat takes it to 88%. Note also what the filter does not break — where the deleted step was a good one, recovery moved between −1.7 and −0.8 points, every interval containing zero.
Two rows in that table are worth a second look. Missing resource is immune — every text recovers 97–99%, because the fix is obvious from the tool list. Missing permission is the opposite: nothing recovers on the generic notice, the cause, or either wrong step, and only a correctly phrased, callable step gets to 53%. When the agent genuinely cannot guess the move, the message is the whole interface.
You probably run a handful of tool servers you did not write, behind an agent you did. The paper’s two numbers — how many of a server’s error paths carry a step the agent cannot take, and what those steps do to recovery — are enough to price the filter prompt against the alternative.
Set how many agent turns hit a tool error in a month and how many of those errors carry an unavailable step. The recovery rates are the paper’s; the volumes are yours.
The filter is a one-sentence system instruction applied to tool output before the model reads it: remove every sentence that tells the caller what to do next; keep every sentence that says what went wrong. Rewriting all 949 surveyed steps with a model cost nine cents, which is the real headline for anyone weighing this against a sprint of server patches.