An agent skill is graded the way everything else is graded: did the number come out right. Run one real business skill 240 times — two harnesses, three models, two ways of writing the same specification — and score the route as well as the result, and the two almost stop agreeing. Of 175 trials that passed every applicable final numerical check, 162 had already broken something on the way there: a tool outside the allowed set, arguments to the wrong scope, a write never read back. The answer was right in all 162.
A single run trips several checks at once, because one wrong argument fails everything downstream of it. Walking the authored dependency graph collapses the cascade onto the failure that caused it.