/p/2026-10-05 · explainer
Paper explainer · 2610.00425 · Vangala, Malik

Functional tests never read the manifest.

A passing test says the code ran on the machine it was built on. It says nothing about whether the dependency list shipped beside it would rebuild it anywhere else — and that list is written by the same model, from the same priors, with nothing checking it. Three coding agents were run over fifty tasks in four languages, each program’s real imports traced at the system-call level, and the declared list compared with what actually loaded. In JavaScript the manifests score 0.22 against the traced truth and in C++ 0.02. Two agents handed the same task agreed on as little as 7% of their packages, and three runs of one agent never once produced the same list.

01 · The problem

Three different answers to “what does this need?”

layer
ecosystem

02 · Scored against what ran

The manifest is accurate in Python and fictional in C++

ecosystem

03 · Run it again

The same agent, the same task, a different list every time

04 · Why nobody noticed

Every agent gets there in the end

show

05 · In your own product

What you install when you accept the manifest illustrative

ecosystem

Results

What the paper actually measured

What it does not show

In practice