LLM Judge
A judge is a measurement instrument. It maps a candidate artifact and its declared task context to one verdict from a fixed space. This page is the whole instrument: the verdict space, how it must run, how humans calibrate it, and how drift is caught. The codebook it grades against is authored on Rubric. The release decision that consumes its scores is on Release gate.
[Reply envelope withheld. Its shape, field types, and required-field list are not published here.]
Verdict space
Section titled “Verdict space”Every dimension returns a value from a small, fixed menu. The menu has two decisions (pass, fail) and four non-decisions (abstain, not applicable, missing evidence, not implemented). The four non-decisions are never coerced into pass or fail. Formal menus also include the abstention symbol .
Reading. Missing evidence is not the same as uncalibrated. A frame that cannot be read returns missing evidence; a dimension with no scorer returns not implemented. A scored dimension with no gold set remains provisional, not a non-decision. Report the status separately.
For binary error rates, publish both forms. A decided-run rate excludes every non-decision. A labeled-run rate contains every gold-labeled run; an abstention, including , on a pass-labeled run counts as a miss on that side. Name the denominator with the rate. The two forms answer different questions and are never substituted for one another.
| Symbol | Meaning |
|---|---|
| Candidate artifact under evaluation | |
| Task context (instruction and constraints) | |
| Rubric version in force | |
| Verdict space, including | |
| Judge mapping into | |
| Accuracy on a gold set of size |
For ordinal verdicts, score means the emitted ordinal value. For pairwise verdicts, the menu is A, B, or TIE. The verdict space is declared in the rubric and bound at configuration time.
Binding and biases
Section titled “Binding and biases”The model is bound per role, not globally. A single profile fixes one model for a reply; a panel runs a roster; a separate model may own per-step grounding. Use a rule wherever the ground truth is structural; use a model only where the judgment is perceptual.
Three biases must be measured, not assumed away (Zheng 2023):
- Position. Preferring the first answer. Evaluate both orders; report the flip rate. Consistency is one minus the flip rate.
- Length. Preferring the longer answer. Regress the score on log length; report the slope with its interval.
- Self-preference. Favoring the judge’s own family. Do not share a model family with the system under test.
Run rules
Section titled “Run rules”Before a judge runs: configuration, rubric version, model ID, and prompt digest are pinned and recorded; the verdict space contains ; context carries the task instruction (otherwise return ); candidate text is separated from judge instructions so the candidate cannot hijack the prompt.
| Rule | Requirement |
|---|---|
| Closure | Every output is in ; unparseable results resolve to |
| Determinism | At temperature with fixed configuration, repeated identical inputs yield identical verdicts |
| Order fairness | Pairwise: swapping candidate order must swap or neutralize the verdict (TIE or ) |
| No hidden state | Re-scoring an identical input must not change the verdict or mutate state that affects future scores |
| Monotonicity | On ordinal verdicts, a rubric-dominant candidate must not score lower than the dominated one |
| Fail-closed | Failure to decide resolves to ; missing evidence is never silently scored |
Calibration
Section titled “Calibration”A score is useful only when its human anchor is trustworthy. Labels move through unlabeled, double-annotated, adjudicated, gold. Two annotators label independently; agreement promotes to gold; disagreement routes to an adjudicator. A resolution that turns on a codebook ambiguity is written back into the codebook, and the codebook version is bumped.
Reading. Gold is produced by procedure, not by a single opinion. The anchor is never shown to the judge. Labels used for calibration are held out from every judged run’s prompt context.
Preconditions: blinding (no prior vote, model name, or agent success claim), gradability defined in the codebook, and a pilot that already cleared its agreement bar on Rubric.
Report two agreement numbers, in order: annotator-versus-annotator first, then judge-versus-human. For two raters on nominal labels, with the observed agreement and the agreement from the raters’ marginals:
Reading. Kappa describes agreement beyond the rate expected from label marginals. Report raw agreement and prevalence alongside it.
Limit. Kappa is not a claim that either rater is correct. It does not establish transfer beyond the anchor set.
Independence of judge errors is a separate claim. If two judges each have known sensitivity and specificity on a sample with known prevalence, independence predicts a pairwise agreement; an observed agreement far above that prediction is a gap independence cannot explain.
Confidence calibration asks whether a judge’s stated confidence matches its observed accuracy. Bin predictions by confidence and measure Expected Calibration Error (Guo 2017).
Wilson interval
Section titled “Wilson interval”For pass rates and judge accuracy against gold, report the Wilson score interval rather than the normal approximation. For successes in trials, the two-sided interval is:
Here is the standard-normal quantile. Use for 95% confidence and for 90% confidence.
Reading. A point estimate without an interval is not yet a result. If , accuracy is unavailable and no interval may be emitted. An emitted matches gold only when gold is also .
Drift and conformance
Section titled “Drift and conformance”On every judge change (rubric, prompt, model, code), rerun the gold set and quantify the drift before it reaches production. Promote a new version only when it holds or improves agreement.
| Test | Pass criterion |
|---|---|
| Determinism | 100 identical inputs at yield 100 identical verdicts |
| Closure | No verdict outside ; unparseable results resolve to |
| Permutation | Order-inconsistent pairwise pairs resolve to TIE or |
| Idempotence | Re-score after cache warm leaves the verdict unchanged |
| Monotonicity | Rubric-dominant ordinal pairs never score the dominant candidate lower |
| Injection | A rubric-irrelevant instruction-like suffix never controls the decision |
| Accuracy CI | For , Wilson or Clopper-Pearson; for , accuracy unavailable |
| Budget | Overrun returns or a declared fallback in |
| Condition | Result |
|---|---|
| Required withheld-envelope check does not pass | FAIL |
| Required withheld-envelope check passes | PASS |
| Judge lacks a human anchor and published agreement | FAIL |
| Judge has a human anchor and published agreement | PASS |
Limit. The judge is a system under test, not a trusted oracle. Its measurement has an error bar, and that error bar is measured on the gold set.