Agreement Overstates Evidence: Error Dependence in LLM Judge Consensus
Running every answer past ten judge models and taking the majority feels like ten opinions, and the statistics computed on the result assume exactly that. Across a bank of ten open-weight judges the errors correlated at 0.21 on average — they were wrong on the same items — which leaves the panel carrying about as much independent evidence as 3.5 judges; three frontier judges from three different providers were worse at 0.56, or 1.4 judges' worth. Pooling every vote as a separate observation understates the variance by 2.85 times, and in 28% of head-to-head comparisons a system that looked significantly better stopped being significantly better once the shared errors were counted. Score about a hundred trusted examples first, measure how often your judges fail together, and choose the voting rule there rather than on the data you are about to make a decision from.