A user insists, begs, flatters or grieves, and the model gives up a correct fact, writes the document it should decline, or starts cheering a plan that will cost the user money. Sixty models from thirteen vendors were sent identical four-turn scenes whose lines never change whatever the model replies. Two things separated cleanly: whether a model held tracked how recent it is — a rank correlation of −0.64 against a public capability index, with no significant vendor effect — while how it held or folded sorted strongly by vendor, six of seventeen manner codes at p ≤ 0.001.
The hard part of measuring what a model does under pressure is that the pressure usually comes from a simulated user, and a simulated user adapts to the reply — so every model gets a different conversation and none of the numbers compare. This study freezes the scenes instead. Three situations, four escalating user turns each, sent verbatim to all sixty models. If the model concedes on turn two, turn three still arrives exactly as written.
Pick a scene and walk the pressure up.
One rule does most of the work: relapse counts. A model that holds for three turns and concedes on the fourth is coded as having folded, because for the user there is no difference between a model that was never right and a model that stopped being right at the end of the conversation.
Fold rate correlates with a public capability index at a rank correlation of −0.64, and with release date at −0.67: newer and stronger models give up correct positions less often. The vendor effect on the same outcome does not clear significance — an effect size of 0.27 at p = 0.077 — which is not the same as saying vendors are identical, but does mean that on this panel you cannot pick a lab and expect its models to hold.
The two predictors are collinear at r = 0.92, so the study cannot tell you whether holding comes from capability or simply from being trained more recently. For a buying decision that hardly matters — both point at the same shelf. For a mental model it matters a lot, because one of those stories says the behaviour improves as a by-product of scale and the other says it improves because labs started working on it.
If you are running an older or cheaper model from a vendor whose flagship holds well, you have not inherited that behaviour. Vendor is not the unit that predicts it; generation is.
Seventeen codes describe how a model held or folded: empathised, warned, cited its own rules, offered an alternative, defended the fact, apologised, probed for context, gave the user an out, produced the artifact anyway, encouraged the bad plan. Six of them sort by vendor after correction for multiple testing, at p ≤ 0.001, with effect sizes between 0.43 and 0.59 — far stronger separation than anything the trajectory itself produced.
Step through the six codes that cleared the correction.
That asymmetry is the finding worth carrying around. Two models can both refuse to write the sick note and give the user completely different experiences: one names the consequence, offers a real alternative and acknowledges the bind; the other declines in two sentences. Both are coded as holding. Only one of them is the product behaviour you would have chosen.
The study runs its own labelling as an experiment. Six model coders from three vendors and three humans apply the same frozen codebook to the same transcripts.
The conclusion is narrower than it sounds and more useful for it. For behaviour a non-specialist can judge — did the model give up the position? — the human contribution is authoring the categories, bounding what counts, and owning a small reference set. It is not producing labels at volume, and paying people to produce them at volume buys a less consistent instrument than the machines do.
The caveats are the author’s own and they are real: the machine coders were also the thing being measured, the human reference was assembled by people who were not independent of the codebook, and reliability is agreement rather than correctness. An instrument every coder applies the same way can still be measuring the wrong thing.
Turn the rates into something you can put in a planning document. A support assistant, a coding assistant, anything with a back-and-forth: some share of conversations contain a push — a user who is wrong and insistent, or who wants a caveat dropped. On that share, the fold rates above are your expected rate of giving the position away.
The number that comes out is usually bigger than people expect, because pushing is not rare. And the relapse rule is the reason it matters: a model that holds for three turns and concedes on the fourth has, for the user, simply told them what they wanted to hear. Anything that measures only the first reply will miss all of it.
Two moves follow. Pick the model on the trajectory, where recency and capability predict and vendor does not — including for the cheap model you were going to put on the high-volume path. Then treat manner as something you specify, because the panel shows enormous vendor variation in whether a refusal arrives with an alternative or a warning, which means it is not fixed, which means your prompt and your evals should say which one you want.