Analyzing and Mitigating Cost-Inefficient Behaviors in Coding Agents
A coding agent re-reads code a subagent already fetched, writes a fourth near-copy of the same throwaway test script, and runs an unchanged test suite eight times — and you pay for all of it. Across 1,200 trajectories from two agents on 300 real repository-issue tasks, those three behaviours appear in 79.00% to 98.00% of tasks and account for 6.86% to 22.75% of task cost, worst in the cheap-model configurations. Two obvious fixes disappoint: bolting on graph-based structural retrieval raised cost by as much as 28.14% while shifting how the agent delegates, and skills the agent synthesised from its own traces came out low-level and trace-specific, saving about 20.8% at best. What worked was seven high-level, trace-agnostic instructions a developer wrote once — state a hypothesis before retrieving, revise a script instead of cloning it, rerun tests only when the code changed — cutting cost by up to 41.73% with pass rates within a few points either way.