Restate a constraint at the moment the agent acts and it obeys. Leave the same constraint back in the history, still in force, and the strongest model tested breaks it in 11.5% of runs — same task, same tools, no attack, nothing adversarial. The paper names that gap and measures it as a matched difference over 412 instances across six tool domains, with histories of about 6,000 tokens spanning 56 to 160 turns. Then it closes the gap with two layers of ordinary engineering: lift each stated rule into a small library when it is said, re-render the matching ones when a task resumes, and put a deterministic check between the model’s proposed call and the tool.