Skip to content

Pipeline

A pipeline is one authoritative flow. Every view on it (the trace workbench, the judge CLI, the reporting dashboard) reads from or writes to the same stages. This page states those stages, the artifact each one emits, and the loop that keeps the judge honest.

An evaluation is not a single score; it is a sequence of stages, each producing a versioned artifact. A run becomes a labeled record only after it passes through all eight. Only the middle of that chain (judge and verdict) is what most people think of as evaluation. The rest is how you know where a number came from.

Reading. A scorecard reads recorded verdicts. It is not a second writer or a replacement for them.

Limit. One writer establishes ownership and provenance. It does not prove that the artifact’s decision is correct.

Eight-stage evaluation pipeline with labeling loop one pipeline, one writer per artifact get data gate score read 1 ingest 2 eval-set 3 assign 4 judge 5 verdict 6 records 7 report 8 decide active labeling disagreement to new judge to new verdict labels never rewrite a past verdict; they mint a new judge
Figure 1.One pipeline. Color marks the job: get data, gate, score, read. Animated edges are emphasis only. The active-labeling loop returns from verdict to judge; it does not rewrite ingest.

Each stage is the only writer of its artifact, and the artifact is immutable once written.

  • Ingest. Raw runs enter with their full context (task, ordered steps, per-step observations). Ingestion stamps provenance and computes a content hash.
  • Eval-set. Runs are pulled into a versioned eval-set: a named collection with a card, a schema, and a manifest of records. The set is the unit of sharing and reproducibility.
  • Assignment. Each item is assigned to one judge profile (single model, panel, or rule), and to one or more human annotators for the gold set. Assignments are versioned.
  • Judge. The judge returns a candidate judgment: per-dimension decisions plus reasons. The judge is bound to a model, a prompt, and a rubric version.
  • Verdict. The verdict stage validates the candidate judgment and writes the immutable per-item verdict. Each verdict carries its provenance (model, prompt, rubric, code revision) and is hashed to an idempotency key.
  • Records. The records stage writes one append-only official record that references that verdict. A rerun creates a new record; nothing is overwritten.
  • Report. A scorecard is read back from the records. It separates judged from total, reports coverage, and shows the nn behind every rate.
  • Decision. A release decision (promote, hold, investigate) is made from the scorecard, the agreement on the gold set, and the drift since the last version.

Human labels are scarce and expensive. The loop spends them where they move a number: mine the judge-human disagreements, label those, improve the judge, re-score. After enough iterations, the loop may improve the judge on the sampled family. Calibrated status still requires a human anchor and published agreement for each dimension.

The loop returns from verdict to judge, not from verdict back to ingest. A label is never retroactively added to a run that was already judged; the new label produces a new judge, which re-scores the same run into a new verdict. The two records share a trace content hash and differ on judge version.

Limit. Active labeling improves the instrument for the family you sampled. It does not transfer, by itself, to a new task family or a new interface surface.

Every artifact in the pipeline has a version. The eval-set version, the rubric version, the judge model, the judge prompt, the judge code, and the taxonomy version are all recorded. A change to any of them is a change to the standard and must be visible in the scorecard.

Reading. Version drift is not a footnote. It is the difference between a regression and a changed ruler.

Limit. A shared gold comparison tells you how two judges relate on that gold set. It does not prove the relation transfers to a new task family or interface surface (JudgeBench 2025).

Artifact lineage and scorecard comparison boundary one record pins the standard that produced it rubric version judge version trace content hash eval set card + records verdict decisions + reasons record append-only scorecard rates + coverage record pin set: eval set, rubric, judge model and prompt, judge code, taxonomy same pin set: compare model releases directly changed judge: compare both on the same gold set first
Figure 2.Artifact lineage for a scored run. The record pins its set, rubric, judge, and taxonomy. Model releases can compare directly only when that pin set matches. A changed judge is checked on shared gold first.

The next pages define the eval-set contract and the rubric the judge is bound to. Classification of failures, the judge itself, and the release decision follow after those.

WorkCitation
JudgeBench (2025). A Benchmark for Evaluating LLM-Based Judges. ICLR.