/p/2026-10-02 · explainer
Paper explainer · 2609.39229 · Onofri and Di Pietro

Ten cheap judges for a sixty-fourth of the price.

Grading answers for faithfulness to their sources with a frontier model costs $2.52 per thousand items and scales with your traffic. Ten small open-weight judges from unrelated families, combined by a plain logistic regression, cost about four cents per thousand and keep a median 93% of the frontier judge’s agreement with the gold labels once chance agreement is discounted. It is not a blanket win — across eight sets the panel clearly beat the frontier judge on one and clearly lost on three — and the transferable part is not the panel but the cheap test that tells you, before you switch, whether yours is one of the cases where it works.

01 · The bill

A judge you call on every answer is a running cost, not a setup cost

02 · The result

Close on most sets, well behind on a few, ahead on one

03 · The test

Two things about your judges decide it, and both are free to check

The admissibility test, run on your own panel

04 · The comparison

It also holds up against detectors built for this one job

compare the panel against

05 · Why I care

Price the switch before you argue about it illustrative

Results

What the paper actually measured

What it does not show

In practice