LLM-as-a-Judge

LLM-as-a-judge is the practice of using a language model to grade another model's output against a rubric, a reference answer or a competing output, in place of a human rater or a fixed metric. It became the default method for scoring open-ended text because it is cheap, fast and can explain its scores, but a judge is itself a model with measurable biases, and its verdicts are only as trustworthy as its agreement with human labels on the task at hand.

How It Works

Three set-ups are common. In pointwise scoring, the judge reads one output and rates it against a rubric, often criterion by criterion. In pairwise comparison, it reads two outputs and picks the better one. In reference-based grading, it checks an output against a known-good answer. G-Eval (Liu et al., March 2023) set the pattern for rubric scoring: the judge is given criteria, generates evaluation steps with chain-of-thought reasoning and then fills in a form. With GPT-4 as the judge it reached a Spearman correlation of 0.514 with human ratings on summarization, which beat earlier automatic metrics and was still far from perfect agreement.

The case for the method was made by Zheng et al. (June 2023), who introduced MT-Bench and reported that strong judges such as GPT-4 matched both controlled and crowdsourced human preferences with over 80% agreement, the same level of agreement they found between humans. That result covers pairwise preference on chat responses with 2023 models; it should not be read as a general guarantee for other tasks or rubrics.

Known Biases

BiasWhat happensSource
PositionThe verdict depends on the order in which candidates are shownWang et al., May 2023; Zheng et al., June 2023
VerbosityLonger answers are favored regardless of qualityZheng et al., June 2023
Self-preferenceThe judge rates its own generations higher than human raters doPanickssery et al., April 2024
LeniencyThe judge tends to mark answers correctThakur et al., June 2024

Position bias is the best documented. Wang et al. showed that a ranking could be "easily hacked" by changing only the order of the responses: with ChatGPT as the evaluator, Vicuna-13B could be made to beat ChatGPT on 66 of 80 test queries. Self-preference has a proposed mechanism. Panickssery, Bowman and Feng found that models including GPT-4 and Llama 2 can distinguish their own outputs from others' at better-than-chance accuracy, and reported a linear correlation between that self-recognition ability and the strength of self-preference. The G-Eval authors had already flagged a related concern, that LLM evaluators may favor LLM-generated text over human writing.

The list is longer than four. Ye et al. (October 2024) catalogued 12 potential biases and built a framework, CALM, to quantify them, concluding that significant biases persist on specific tasks even in advanced models. Most of these studies used models from 2023 and 2024. Whether newer judges show the same biases at the same magnitude is an empirical question that the papers cited here do not answer, so the safe assumption is that each bias has to be tested for rather than presumed fixed.

Calibration Against Human Labels

A judge is a measuring instrument and has to be validated like one. The standard procedure is to have people label a sample, run the judge on the same sample, and compare. Thakur et al., who tested thirteen judge models, found that only the largest aligned reasonably with humans and that even those remained well below inter-human agreement. They also showed that the choice of statistic matters: judges with high percent agreement can still assign very different scores, differing from human-assigned scores by up to 5 points, which is why chance-corrected measures are more informative than raw agreement. The authors describe their setting as a simplified one, so harder rubrics are unlikely to fare better.

Several mitigations are documented. Wang et al. proposed having the judge write out its evidence before scoring, averaging over both orderings of a pair, and sending the hardest cases to a person. Anthropic's guide to agent evaluation (January 2026) recommends that LLM-as-judge graders be closely calibrated with human experts and that teams read transcripts and grades regularly to confirm the grader measures what was intended. Two further practices follow from the bias findings: use a judge from a different model family than the system being graded, and write rubrics as explicit, separately gradeable criteria. Calibration is not permanent. A change to the judge model, the rubric or the kind of output being graded calls for a fresh comparison against human labels.

When Programmatic Checks Are Better

If correctness can be computed, it should be. Tests that pass or fail, a database row that exists or does not, a schema that validates, a number within tolerance: these checks are deterministic, cost almost nothing and cannot be talked into a higher score. Anthropic's guide describes code-based graders as fast, cheap, objective and reproducible, and model-based graders as flexible but non-deterministic and more expensive.

Judges earn their place where no such check exists: tone, helpfulness, faithfulness of a summary, whether an explanation answers the question asked. In agent evals the two are usually combined, with a programmatic verifier for the final state and a rubric judge for qualities that cannot be asserted in code. A judge's score can also become a training signal, in reinforcement fine-tuning or prompt optimization. That raises the stakes on every bias above, because an optimizer will find and exploit whatever the judge over-rewards — the dynamic described under reward hacking.

Further Reading