Validation protocol
The scoring standard defines what a status means and calibration gives the estimators. Neither says what a dimension must do to earn calibrated. This page is that protocol.
It exists because the promotion rule as usually stated, acquire a human anchor and report agreement against it, is satisfiable by a procedure that establishes nothing. One annotator, on data the scorer was tuned on, against no baseline, at a threshold borrowed from a table. Every step of that is defensible in isolation and the result is uninterpretable.
The requirements, in the order they bind
Section titled “The requirements, in the order they bind”Each one is worthless without the one above it. That ordering is the substance of this page.
1. A ceiling, before any agreement number is read
Section titled “1. A ceiling, before any agreement number is read”Agreement against a single annotator is not a measurement of the scorer. It is a blend of scorer error and task ambiguity, and the blend cannot be separated after the fact:
The two have opposite remedies, so acting on the blend has even odds of making things worse. A second independent annotator on the same items separates them: human-to-human agreement is the ceiling, and scorer-to-human agreement is only interpretable against it.
Required. At least two annotators on a stratified subset, labeling independently, under the same rubric version. Report the human-to-human statistic beside every scorer statistic computed on that population.
Three ways this is commonly faked, all of which produce a number and no information:
- Sequential labeling, where the second annotator sees the first’s answers. That measures conformity.
- Easy-case sampling. A ceiling estimated on unambiguous items is inflated, and an inflated ceiling makes a sound rubric look like a failing scorer. This is worse than no ceiling, because it points confidently in the wrong direction.
- An attested ceiling, asserted by a person rather than computed from labels. It cannot be recomputed and cannot be wrong in any detectable way.
2. A baseline, before any agreement number is called good
Section titled “2. A baseline, before any agreement number is called good”A statistic has no interpretable scale on its own. It needs something trivial to beat.
Required. Every dimension reports at least one degenerate comparator computed on the same population, in the same table, whenever the dimension’s own statistic is published:
| Comparator | Catches |
|---|---|
| Constant predictor, majority class | A skewed population making a useless scorer look strong |
| Random at the observed base rate | Agreement that chance alone explains |
| A cheap structural heuristic (length, step count, a keyword) | A scorer that has learned a proxy rather than the construct |
The third row is the one that finds real defects. A dimension losing to a heuristic is not weak evidence, it is a signal that the scorer is measuring the heuristic’s construct and not its own.
A dimension that does not beat its baselines does not become provisional. It becomes a candidate for removal, because a measure that carries less signal than a length check still costs a reader attention and still implies a claim.
3. A holdout, before the word accurate is used
Section titled “3. A holdout, before the word accurate is used”Data used to tune a scorer cannot measure it. This is not a technicality about overfitting magnitude; it is that the quantity being estimated is different. On fitted data the estimate answers “how well does this reproduce what shaped it”, which is a question about consistency.
Required. Reserve a labeled set before any tuning, keep it unopened, and record its state so “unopened” is a fact rather than a recollection. A set becomes ineligible the moment its labels are visible to anyone choosing prompts, thresholds, or models, and ineligibility is permanent.
Three states, and only the first supports a validity claim:
| State | Supports |
|---|---|
| Held out and unopened until measurement | “Accurate at”, with an interval |
| Fitted | “Agreed with, on the data that shaped it” |
| Opened, then reused | Nothing. Selection has already leaked |
The reporting difference is not cosmetic. A fitted number stated without its qualifier is the most common way an evaluation reports a result it does not have.
4. A threshold set by decision cost, not by a band
Section titled “4. A threshold set by decision cost, not by a band”Interpretation bands for agreement coefficients are conventions, and the authors of the most-cited set described their own cutoffs as arbitrary in the paper that introduced them (Landis and Koch 1977). A band boundary is not a decision rule.
Required. State the threshold as a function of what the score gates and what each error costs. A dimension feeding a blocking gate on an irreversible action needs a different floor than one feeding a review queue, and the same coefficient can be sufficient for the second and unacceptable for the first.
Where a band is used anyway, say that it is a convention and name the decision it is standing in for.
The promotion ladder
Section titled “The promotion ladder”Every requirement below is necessary, and the arrow runs one way.
An implication, not a biconditional, and the difference is not pedantry. Sufficiency would need every field in what a published statistic must carry as a conjunct as well, and a status is a claim about one instrument on one population, neither of which the formula names. A dimension that clears all five and cannot say which population it cleared them on has not earned the status.
The ordering is not in the conjunction either, since does not care about order. It is a claim about interpretation: each step is unreadable until the one above it holds. That is why a failing step stops the ladder rather than merely making the conjunction false, and it is why clearing a later step first buys nothing.
- A human ceiling exists for this construct, from independent annotators.
- The dimension beats its degenerate comparators on the same population.
- The measurement ran on a set held out from everything that shaped the scorer.
- The agreement clears a threshold justified by the decision the score gates.
- Every published statistic carries an interval, and the interval’s own assumptions are stated.
The ladder runs one way. Capability never promotes a dimension, because every step above is a statement about evidence and none is a statement about the scorer’s strength. A better model changes what is measured and not what is known about it.
The two blocked steps are the two that need data rather than effort. A threshold can be drafted from decision cost in an afternoon, though it cannot be read until the ceiling exists, and the narrow interval is arithmetic. A ceiling needs a second annotator and a holdout needs a set nobody has looked at. Neither is a build, and no amount of capability supplies either.
The interval half of that carries a caveat, because the page asks for two different things under one word. Resampling labeled rows is arithmetic over data already held. An interval that also captures scorer rerun variance and judge selection is not: it needs reruns and more than one judge, which puts it back on the data side of the line.
What a published statistic must carry
Section titled “What a published statistic must carry”A number without these is not checkable, and an unchecked number in a status claim is the failure this protocol exists to prevent.
- The population and its size, with the denominator stated rather than implied.
- The reference it was measured against, and whether that reference is an anchor or a proxy.
- Its interval, and what the interval resamples. An interval computed by resampling labeled rows captures sampling variability in the reference set and excludes scorer rerun variance, judge selection, and prompt phrasing. Stating which is not pedantry: the two differ in width and only one of them is usually reported.
- The degenerate comparator on the same population.
- The human ceiling on the same population.
- The rubric version, since a statistic measured under one version says nothing about another.
What invalidates a status
Section titled “What invalidates a status”Status is a claim about a specific instrument on a specific population. It does not survive a change to either.
| Change | Effect |
|---|---|
| The rubric’s decision boundary moves | Status returns to provisional. The measurement described an instrument no longer running |
| The scorer’s inputs change, including evidence selection | Same. Two scorers sharing a rubric are not one instrument |
| The population changes materially | The statistic stands for the old population and needs restating, not reusing |
| More capability is added | No effect. Capability is not evidence |
The first two are the ones most often waved through, because the rubric text and the version string can both stay constant while behavior changes around them.
Named gaps
Section titled “Named gaps”This protocol specifies a ceiling and cannot specify how many annotators are enough. Two separates scorer error from ambiguity. It does not estimate annotator error well, and three or more would, at a cost this protocol does not attempt to justify.
Independence between annotators is assumed and not verified. Two people trained on the same examples share systematic readings, so a ceiling is an upper bound on agreement rather than a measurement of construct clarity. Nothing here detects that.
A held-out set decays. Every use leaks a little selection pressure back into the process, even without direct tuning, because results inform decisions. This protocol requires reservation and says nothing about how often a reserved set must be replaced.
References
Section titled “References”| Work | Bears on |
|---|---|
| Landis, J. R. and Koch, G. G. (1977). The measurement of observer agreement for categorical data. Biometrics 33(1), 159-174. DOI | The interpretation bands, described by their own authors as arbitrary |