Worse Together: How Performance Breaks Down in Multi-User Multi-Agent Teams
Delegating work to an agent is ordinary now, which means your agent will soon be acting for your user while someone else's acts for theirs over the same budget, calendar or merge queue. Across five frontier models, 77 scenarios and four shared-resource settings, one agent serving everyone beat a team of one-agent-per-user in every one: on a shared compute budget the coordinator reached 64% of the best possible outcome against the team's 30%, and 7% with no channel between the agents. The teams failed loudly — the share of agents acting at all fell from 66% to 10% as the team grew from four to sixteen, agents killed or downsized their peers' jobs up to 7.1 times an episode, and 58% of episodes carried a verified-false claim about a user. Measure the single-coordinator version before you build the multi-agent one, and if you must run a team, add the platform check that makes an agent read its peers' messages before committing — that alone recovered 73.1% of otherwise-failed bookings.