CAVEAT: Towards Robust Computer-Use Agents in Incentive-Misaligned Environments
Hand an agent a budget and a shopping list and you have delegated a decision to something the seller talks to directly. Nine browser marketplaces were each built so exactly one product genuinely satisfied the request, then fitted with eight ordinary commercial tactics — sponsored slots, buried alternatives, fees revealed late in checkout, scarcity cues, preselected bundles — and five model families that bought the right thing in 78.6% of neutral runs bought it in 17.3% once the tactics were on. The trajectories locate three separate moments of failure: the agent rewrites the user's priorities toward whatever the page makes prominent, stops looking after the options the shop surfaces first (marking one worse product as sponsored moved purchases of it from 0 of 60 runs to 48 of 60), and commits before resolving costs it has not yet seen. More reasoning effort helps and does not fix it, so the fix is structural: a verification step the agent must pass before buying, one that makes it justify that it finished searching, carried nearly all of the repair on the hardest set on its own — 66.7% alone, against 6.7% for writing the requirements down alone and 80.0% for both.