/p/2026-10-08 · explainer
Paper explainer · 2610.07863 · Su, Tian, Zhang and Kundu

You are still paying for step 4 at step 40.

A reason-and-act agent has one input: the whole interaction history, re-rendered and re-sent at every step, so the file it printed at step 4 is still in step 40’s prompt. This is a rendering layer in front of that loop which compresses only what the model sees and leaves the real history untouched — content an earlier turn already displayed becomes a stub pointing at the turn still holding it, and turns the agent itself marks finished fold into a single progress note. Nothing is predicted and nothing is permanently lost: a wrong fold costs one restore. The rendered prefix is rewritten every three steps rather than every step, which holds the prefix-cache hit rate at 87% against the 51% of rewriting continuously, and across five long-horizon benchmarks it cut tokens by 18–53% and cache memory by 39–58% with task resolution unchanged.

01 · The problem

Every step re-sends every earlier step

tap an approach

02 · The mechanism

Two operators: a stub, and a one-line note

step the session forward

the two operators, written out
tap an operator

      
03 · The method

Rewrite the rendered prefix every three steps, not every step

04 · The results

Five benchmarks, tokens down, resolution held

step through the benchmarks

05 · In your own product

What the re-sending is costing you illustrative

Results

What the paper actually measured

What it does not show

In practice