Skip to content

LLM Judge

A judge is a measurement instrument. It maps a candidate artifact and its declared task context to one verdict from a fixed space. This page is the whole instrument: the verdict space, how it must run, how humans calibrate it, and how drift is caught. The codebook it grades against is authored on Rubric. The release decision that consumes its scores is on Release gate.

[Reply envelope withheld. Its shape, field types, and required-field list are not published here.]

Every dimension returns a value from a small, fixed menu. The menu has two decisions (pass, fail) and four non-decisions (abstain, not applicable, missing evidence, not implemented). The four non-decisions are never coerced into pass or fail. Formal menus also include the abstention symbol ⊥\bot.

Closed verdict menu with decisions and non-decisions decisions pass fail non-decisions abstain no parse n/a out of missing evidence not impl. non-decisions never become pass or fail; coverage reports how often the instrument decided
Figure 1.Closed menu. Decisions and non-decisions stay separate. Coverage is the share of dimensions the instrument actually decided.

Reading. Missing evidence is not the same as uncalibrated. A frame that cannot be read returns missing evidence; a dimension with no scorer returns not implemented. A scored dimension with no gold set remains provisional, not a non-decision. Report the status separately.

For binary error rates, publish both forms. A decided-run rate excludes every non-decision. A labeled-run rate contains every gold-labeled run; an abstention, including ⊥\bot, on a pass-labeled run counts as a miss on that side. Name the denominator with the rate. The two forms answer different questions and are never substituted for one another.

candidate + context artifact, task, rubric Judge J fixed decision rule declared verdict V one verdict or abstention human labels held-out anchor calibration labels do not enter J
Figure 2.Measurement boundary. Candidate and context enter the judge. Labels enter calibration only. Animated edges are emphasis only.
SymbolMeaning
xxCandidate artifact under evaluation
ccTask context (instruction and constraints)
RRRubric version in force
VVVerdict space, including ⊥\bot
JJJudge mapping into VV
p^\hat{p}Accuracy X/mX/m on a gold set of size mm

For ordinal verdicts, score means the emitted ordinal value. For pairwise verdicts, the menu is A, B, or TIE. The verdict space is declared in the rubric and bound at configuration time.

The model is bound per role, not globally. A single profile fixes one model for a reply; a panel runs a roster; a separate model may own per-step grounding. Use a rule wherever the ground truth is structural; use a model only where the judgment is perceptual.

Three biases must be measured, not assumed away (Zheng 2023):

  • Position. Preferring the first answer. Evaluate both orders; report the flip rate. Consistency is one minus the flip rate.
  • Length. Preferring the longer answer. Regress the score on log length; report the slope with its interval.
  • Self-preference. Favoring the judge’s own family. Do not share a model family with the system under test.

Before a judge runs: configuration, rubric version, model ID, and prompt digest are pinned and recorded; the verdict space contains ⊥\bot; context carries the task instruction (otherwise return ⊥\bot); candidate text is separated from judge instructions so the candidate cannot hijack the prompt.

RuleRequirement
ClosureEvery output is in VV; unparseable results resolve to ⊥\bot
DeterminismAt temperature 00 with fixed configuration, repeated identical inputs yield identical verdicts
Order fairnessPairwise: swapping candidate order must swap or neutralize the verdict (TIE or ⊥\bot)
No hidden stateRe-scoring an identical input must not change the verdict or mutate state that affects future scores
MonotonicityOn ordinal verdicts, a rubric-dominant candidate must not score lower than the dominated one
Fail-closedFailure to decide resolves to ⊥\bot; missing evidence is never silently scored

A score is useful only when its human anchor is trustworthy. Labels move through unlabeled, double-annotated, adjudicated, gold. Two annotators label independently; agreement promotes to gold; disagreement routes to an adjudicator. A resolution that turns on a codebook ambiguity is written back into the codebook, and the codebook version is bumped.

Reading. Gold is produced by procedure, not by a single opinion. The anchor is never shown to the judge. Labels used for calibration are held out from every judged run’s prompt context.

Preconditions: blinding (no prior vote, model name, or agent success claim), gradability defined in the codebook, and a pilot that already cleared its agreement bar on Rubric.

Report two agreement numbers, in order: annotator-versus-annotator first, then judge-versus-human. For two raters on nominal labels, with pop_o the observed agreement and pep_e the agreement from the raters’ marginals:

κ=po−pe1−pe\kappa = \frac{p_o - p_e}{1 - p_e}

Reading. Kappa describes agreement beyond the rate expected from label marginals. Report raw agreement and prevalence alongside it.

Limit. Kappa is not a claim that either rater is correct. It does not establish transfer beyond the anchor set.

Illustrative Cohen kappa from a two-by-two table rater 2 pass fail rater 1 pass fail 40 10 10 40 p_o = (a+d)/n = 0.80 p_e from marginals = 0.50 kappa = 0.60 illustrative n = 100; report p_o and prevalence beside kappa when p_e = 1, kappa is undefined: drop the replicate, do not clamp
Figure 3.Illustrative two-by-two. Kappa is computed from the stated cell counts inside the figure. Teaching numbers only.

Independence of judge errors is a separate claim. If two judges each have known sensitivity and specificity on a sample with known prevalence, independence predicts a pairwise agreement; an observed agreement far above that prediction is a gap independence cannot explain.

Observed judge agreement against what independence predicts predicted, if judge errors were independent 0.72 observed 0.95 agreement independence cannot explain 0.5 0.6 0.7 0.8 0.9 1.0 pairwise agreement
Figure 4.Illustrative test of judge independence. Predicted agreement is computed in the figure from sensitivity 0.8, specificity 0.9, and prevalence 0.75. Observed 0.95 is a teaching contrast, not a measured study.

Confidence calibration asks whether a judge’s stated confidence matches its observed accuracy. Bin predictions by confidence and measure Expected Calibration Error (Guo 2017).

For pass rates and judge accuracy against gold, report the Wilson score interval rather than the normal approximation. For p^=X/m\hat{p} = X/m successes in mm trials, the two-sided 1−α1-\alpha interval is:

p^+z22m±zp^(1−p^)m+z24m21+z2m\frac{\hat{p} + \frac{z^2}{2m} \pm z \sqrt{\frac{\hat{p}(1-\hat{p})}{m} + \frac{z^2}{4m^2}}}{1 + \frac{z^2}{m}}

Here z=Φ−1(1−α/2)z = \Phi^{-1}(1-\alpha/2) is the standard-normal quantile. Use z=1.96z = 1.96 for 95% confidence and z=1.645z = 1.645 for 90% confidence.

Reading. A point estimate without an interval is not yet a result. If m=0m = 0, accuracy is unavailable and no interval may be emitted. An emitted ⊥\bot matches gold only when gold is also ⊥\bot.

Illustrative Wilson score interval for a binomial proportion Wilson interval (illustrative) 0.00 0.25 0.50 0.75 1.00 0.551 0.880 p-hat = 0.750 k = 18, n = 24, z for 95%; bounds computed in the figure a point estimate without an interval is not yet a result
Figure 5.Illustrative Wilson interval for 18 successes in 24 trials at 95%. Bounds are computed in the figure from those inputs. Teaching counts only.

On every judge change (rubric, prompt, model, code), rerun the gold set and quantify the drift before it reaches production. Promote a new version only when it holds or improves agreement.

TestPass criterion
Determinism100 identical inputs at T=0T=0 yield 100 identical verdicts
ClosureNo verdict outside VV; unparseable results resolve to ⊥\bot
PermutationOrder-inconsistent pairwise pairs resolve to TIE or ⊥\bot
IdempotenceRe-score after cache warm leaves the verdict unchanged
MonotonicityRubric-dominant ordinal pairs never score the dominant candidate lower
InjectionA rubric-irrelevant instruction-like suffix never controls the decision
Accuracy CIFor m≥1m \ge 1, Wilson or Clopper-Pearson; for m=0m=0, accuracy unavailable
BudgetOverrun returns ⊥\bot or a declared fallback in VV
ConditionResult
Required withheld-envelope check does not passFAIL
Required withheld-envelope check passesPASS
Judge lacks a human anchor and published agreementFAIL
Judge has a human anchor and published agreementPASS

Limit. The judge is a system under test, not a trusted oracle. Its measurement has an error bar, and that error bar is measured on the gold set.

WorkCitation
Cohen, J. (1960). A coefficient of agreement for nominal scales. Educational and Psychological Measurement 20(1), 37-46.
Krippendorff, K. (1970). Bivariate agreement coefficients for reliability of data. Sociological Methodology 2, 139-150.
Feinstein, A. R. and Cicchetti, D. V. (1990). High agreement but low kappa: I. The problems of two paradoxes. Journal of Clinical Epidemiology 43(6), 543-549.
Guo, C. et al. (2017). On calibration of modern neural networks. ICML.
Zheng, L. et al. (2023). Judging LLM-as-a-judge with MT-Bench and Chatbot Arena. arXiv:2306.05685.
Brown et al. (2001). Interval Estimation for a Binomial Proportion. Statistical Science.