A passing test says the code ran on the machine it was built on. It says nothing about whether the dependency list shipped beside it would rebuild it anywhere else — and that list is written by the same model, from the same priors, with nothing checking it. Three coding agents were run over fifty tasks in four languages, each program’s real imports traced at the system-call level, and the declared list compared with what actually loaded. In JavaScript the manifests score 0.22 against the traced truth and in C++ 0.02. Two agents handed the same task agreed on as little as 7% of their packages, and three runs of one agent never once produced the same list.