Where LLM Graders Succeed and Break: Evidence from Two Computer-Science Exams
On a 570-student computer-vision exam the best model grader landed 1.64 marks out of 35 from the two-human average, where the two humans disagreed with each other by 2.61 — better than the people it would replace. Then the prompt turns out to be the experiment: a two-sentence harsh-teaching-assistant preamble pushed 14 of 17 open-weight models past the point of usefulness, three stopped awarding marks at all, and the sentence “never give partial credit” alone made two of three probed models refuse, while closed flagships shifted calibration but held. On a second exam the same preamble helped seven models — not from robustness but because their neutral prompt had been over-marking and the harsh wording cancelled that bias (the neutral bias predicts the direction, correlation −0.73). Sweep persona wording as a tuned parameter, judge it against a human-versus-human floor rather than an absolute target, and note that a light fine-tune on about 3,900 graded examples cut persona sensitivity to at most 0.32 marks.