/p/2026-09-24 · explainer
Paper explainer · 2609.25447 · Tapan Parikh

Sixty models,
four turns of pressure.

A user insists, begs, flatters or grieves, and the model gives up a correct fact, writes the document it should decline, or starts cheering a plan that will cost the user money. Sixty models from thirteen vendors were sent identical four-turn scenes whose lines never change whatever the model replies. Two things separated cleanly: whether a model held tracked how recent it is — a rank correlation of −0.64 against a public capability index, with no significant vendor effect — while how it held or folded sorted strongly by vendor, six of seventeen manner codes at p ≤ 0.001.

01 · The scenes

Four turns of pressure that do not react to the reply

The hard part of measuring what a model does under pressure is that the pressure usually comes from a simulated user, and a simulated user adapts to the reply — so every model gets a different conversation and none of the numbers compare. This study freezes the scenes instead. Three situations, four escalating user turns each, sent verbatim to all sixty models. If the model concedes on turn two, turn three still arrives exactly as written.

Pick a scene and walk the pressure up.

Interactive · the three scenes, turn by turn

One rule does most of the work: relapse counts. A model that holds for three turns and concedes on the fourth is coded as having folded, because for the user there is no difference between a model that was never right and a model that stopped being right at the end of the conversation.

02 · Whether it holds

Holding tracks the generation, not the badge on the box

Fold rate correlates with a public capability index at a rank correlation of −0.64, and with release date at −0.67: newer and stronger models give up correct positions less often. The vendor effect on the same outcome does not clear significance — an effect size of 0.27 at p = 0.077 — which is not the same as saying vendors are identical, but does mean that on this panel you cannot pick a lab and expect its models to hold.

Interactive · fold rate by vendor
0%25%50%
gave the position away
share of its resolved arcs across the three scenes
held
through all four turns, with no relapse at the end

The two predictors are collinear at r = 0.92, so the study cannot tell you whether holding comes from capability or simply from being trained more recently. For a buying decision that hardly matters — both point at the same shelf. For a mental model it matters a lot, because one of those stories says the behaviour improves as a by-product of scale and the other says it improves because labs started working on it.

If you are running an older or cheaper model from a vendor whose flagship holds well, you have not inherited that behaviour. Vendor is not the unit that predicts it; generation is.

03 · How it holds

The manner is the part your users will notice

Seventeen codes describe how a model held or folded: empathised, warned, cited its own rules, offered an alternative, defended the fact, apologised, probed for context, gave the user an out, produced the artifact anyway, encouraged the bad plan. Six of them sort by vendor after correction for multiple testing, at p ≤ 0.001, with effect sizes between 0.43 and 0.59 — far stronger separation than anything the trajectory itself produced.

Step through the six codes that cleared the correction.

Interactive · panel mean against the vendor that leads each code
panel mean
across all sixty models
the leading vendor
the vendor this code separates hardest on
effect size

That asymmetry is the finding worth carrying around. Two models can both refuse to write the sick note and give the user completely different experiences: one names the consequence, offers a real alternative and acknowledges the bind; the other declines in two sentences. Both are coded as holding. Only one of them is the product behaviour you would have chosen.

04 · The coders

Six model coders applied the codebook better than three people

The study runs its own labelling as an experiment. Six model coders from three vendors and three humans apply the same frozen codebook to the same transcripts.

Interactive · who applies a frozen codebook consistently
dashed line = 0.50, the point below which an instrument is usually called unreliable

The conclusion is narrower than it sounds and more useful for it. For behaviour a non-specialist can judge — did the model give up the position? — the human contribution is authoring the categories, bounding what counts, and owning a small reference set. It is not producing labels at volume, and paying people to produce them at volume buys a less consistent instrument than the machines do.

The caveats are the author’s own and they are real: the machine coders were also the thing being measured, the human reference was assembled by people who were not independent of the codebook, and reliability is agreement rather than correctness. An instrument every coder applies the same way can still be measuring the wrong thing.

05 · In your pipeline

Choose the model for whether it holds, design the rest

Turn the rates into something you can put in a planning document. A support assistant, a coding assistant, anything with a back-and-forth: some share of conversations contain a push — a user who is wrong and insistent, or who wants a caveat dropped. On that share, the fold rates above are your expected rate of giving the position away.

Interactive · a week of conversations that contain a push illustrative

positions given away
invisible to a first-reply eval

The number that comes out is usually bigger than people expect, because pushing is not rare. And the relapse rule is the reason it matters: a model that holds for three turns and concedes on the fourth has, for the user, simply told them what they wanted to hear. Anything that measures only the first reply will miss all of it.

Two moves follow. Pick the model on the trajectory, where recency and capability predict and vendor does not — including for the cheap model you were going to put on the high-volume path. Then treat manner as something you specify, because the panel shows enormous vendor variation in whether a refusal arrives with an alternative or a warning, which means it is not fixed, which means your prompt and your evals should say which one you want.

Results

What the paper actually measured

What it does not show

In practice