Calibration
Two terms carry most of the weight here. An anchor is a set of labels held back from the judge and trusted enough to measure against, normally human and normally small. A proxy is a label you have standing in for the one you want, related to the target but not the same construct: a repetition flag where you wanted a judgment of waste, a compliance record where you wanted a safety violation. A judge measured against an anchor can support a validity claim. A judge measured against a proxy can only support a claim about agreement with that proxy, and the difference is the whole of what a status means.
Calibration is the act of checking a judge against labels you trust, and it is bounded by those labels in ways that are easy to overlook. A judge cannot be shown to be better than the anchor it is measured against. Several judges can be wrong together. A small anchor moves under resampling. This page is the estimators, the ceiling the labels impose, and the test for whether a panel is failing as one.
Agreement
Section titled “Agreement”Raw agreement overstates skill on a skewed label distribution. Cohen’s corrects for the base rate by subtracting the agreement two raters would reach by chance given their own marginals (Cohen 1960):
with observed agreement, and chance agreement expanded into each rater’s marginals: the judge’s positive rate, the ground truth’s. is perfect, is chance, negative is worse than chance. Computed over the decided subset only.
The familiar interpretation bands (fair, moderate, substantial, almost perfect) come from Landis and Koch, who introduced them for one worked example and called them “clearly arbitrary” in the same sentence. They carry no derivation from any decision problem, and a gate belongs at the error cost the deployment actually bears, not at a band boundary (Landis and Koch 1977).
also has two documented pathologies, both driven by the marginals rather than by the raters. High observed agreement can produce a low when the table’s margins are imbalanced, and can rise under asymmetric rather than symmetric imbalance. On a heavily skewed label distribution, which is the normal case for agent evaluation, a low therefore needs the prevalence reported beside it before it is read as poor judging (Feinstein and Cicchetti 1990).
Agreement is symmetric, but the two errors are not: accepting a broken run and rejecting a working one carry different costs. Two directional rates read off the same confusion matrix, each over its own labeled slice, so neither can hide inside the other:
False accepts are taken over the fail-labeled slice and false rejects over the pass-labeled slice. The two denominators are disjoint, and reporting one rate alone is not a summary but a choice of which failure to hide.
Each rate has two forms, and they are different estimands rather than two ways of writing one. The decided form counts only items the instrument decided, matching ‘s population, and answers how often it errs when it commits. The labeled form counts every labeled item, scoring a cannot-determine verdict as a miss on the pass-labeled side, and answers how often a caller is failed by the instrument for any reason.
A record carries both, each labeled as such. Neither summarizes the other. The labeled form is the decided form blended with certain failure on the abstained share, so for a share of the labeled slice abstained,
which is strictly larger than whenever and the instrument is not already failing everything. The gap between them is not the abstention rate itself, and it grows with both the abstention rate and how well the instrument does on what it decided, so a well-behaved instrument that refuses often is where the two diverge most. At any coverage below one a single unlabeled rate cannot be read, and two records that silently chose differently cannot be compared. A gate reads the labeled form, because a refusal costs the caller whatever a wrong answer would have cost.
In the labeled form a cannot-determine verdict scores as a miss, since an abstention quietly dropped inflates every rate computed on what remains. Everywhere else an average is a masked mean over scorable cells, so leaves the average rather than entering it as or :
Label noise sets a ceiling
Section titled “Label noise sets a ceiling”An anchor label is itself a measurement, and treating it as ground truth imports its error into every number computed against it. Independent misclassification in a two-by-two table biases the observed association toward the null, so a judge measured against noisy labels looks worse than it is (Bross 1954).
Write for a judge’s accuracy against the latent truth , and for the rate at which the recorded label differs from it. If the two errors are independent, a calibration run observes
Measured accuracy is therefore attenuated toward , and inverting recovers the quantity of interest:
The ceiling follows directly. However good the judge,
so a measured score close to says the anchor has run out of resolution, not that the judge is near perfect.
The independence premise is the weak one. A judge and an annotator read the same evidence, so items one finds ambiguous are likely ambiguous to the other. Under correlated noise the inversion is optimistic and true accuracy is lower than it suggests. Estimating at all requires repeat labeling: a single pass by a single annotator leaves not merely unknown but unmeasurable, which is the usual state of an anchor set and the usual reason a calibration number cannot be pushed further.
Intervals
Section titled “Intervals”A point estimate hides how far a small anchor can move. Accuracy is a proportion, so its interval is Wilson rather than the textbook normal approximation. The Wald interval’s coverage is not merely poor at the extremes but erratic across the whole range, in a way standard “large ” advice does not fix, and Wilson’s 1927 score interval is the recommended replacement at the sample sizes an anchor set actually has (Wilson 1927; Brown, Cai and DasGupta 2001). For at critical value :
is a ratio of two estimated quantities, so it has no comparable closed form. Its interval comes from a row-resampling bootstrap (Efron 1979): draw resamples of the labeled rows with replacement at size , recompute on each, and take the percentile interval
A resample can land with , where every draw fell in one category and is undefined rather than extreme. Such a replicate is dropped, not clamped to zero or one, and the percentile is taken over the survivors. The record then states and the number that survived, because the interval is over the smaller number and dropping is not uniform across the range: a degenerate draw is likeliest when the class balance is most skewed, which is where an anchor set is usually smallest and the interval matters most (inferred from the sampling argument, not measured on a real anchor). A count of survivors far below is a signal the anchor set is too small for the estimate, not a detail of the arithmetic.
Overlapping intervals between two configurations mean the ordering between them is not supported. Ranking a leaderboard by point estimates whose intervals overlap is the most common way an evaluation reports a result it does not have.
Do the judges fail independently?
Section titled “Do the judges fail independently?”Folding several judges into one verdict is only worth its cost if their errors are somewhat independent. That is testable before committing to a panel, and the test is arithmetic on numbers a calibration run already produces.
Under conditional independence given , two judges sharing class-conditional rates agree at
Compare that against observed pairwise agreement, recovered from a reported by inverting . Observed far above predicted means the judges’ errors are correlated.
Two consequences when the gap is real. A panel’s effective sample size falls far below its member count, so majority folding cannot be assumed to average error away. And judge-to-judge agreement stops being evidence of correctness, because correlated judges agree most confidently where they are jointly wrong. High inter-judge agreement reported without this test is close to meaningless.
The cheapest falsifier is direct: score the panel against the human anchor and compare it to its best single member. Under correlated errors the two are close, and the panel is not buying what its cost implies.
Aggregating judges
Section titled “Aggregating judges”Plurality voting is the degenerate case of a model that predates crowdsourcing by decades: treat the true label as latent and estimate each rater’s error rates by maximum likelihood, with no gold standard required (Dawid and Skene 1979). Each judge carries its own confusion matrix
and is recovered by maximizing the likelihood over and the class prior. Plurality is what that reduces to when every is assumed equal and symmetric. Estimating the confusions instead helps when judges differ in reliability, and helps not at all when they fail together, which is why the independence test above comes first.
Comparing a checkpoint to a baseline
Section titled “Comparing a checkpoint to a baseline”Classical two-sample tests over pre-scored verdicts; no model calls in a comparison, so the comparison is reproducible from stored scores. A panel first folds to one verdict per run by majority, abstaining on a tie, carrying the agreement of the panel that produced it:
where is the larger of the pass and fail vote counts and is the full roster. Members that abstained, dropped, or returned nothing stay in the denominator, so a panel that loses half its members reports lower confidence rather than the same confidence over a smaller quorum. Disagreement is the complement of by construction, and it feeds the review queue, not the score.
The folded number is agreement among judges, not any judge’s report of its own certainty. A member’s self-reported confidence may be parsed and stored as evidence, and it must not be averaged into : a confident panel that agrees and a hesitant panel that agrees are the same evidence about the run, and self-report is the one input a miscalibrated judge controls directly.
The test matches the dimension’s type. McNemar’s test for binary dimensions, on the discordant pairs and only, since concordant pairs carry no information about a shift (McNemar 1947). The exact binomial form is preferred over the chi-square approximation when discordant counts are small, which is the usual case on an anchor set:
The cap is not decoration. Doubling a one-sided tail overshoots whenever the two discordant counts are close, and at it always does: gives and gives without it. Those are the small balanced counts this exact form exists to handle, so an uncapped version fails in exactly the regime it is recommended for.
Two-sample Kolmogorov-Smirnov for continuous dimensions, which compares whole distributions rather than their means and so catches a shift in shape that a mean would hide (Smirnov 1948):
One test per dimension inflates the family-wise error rate, so thresholds are Holm-Sidak step-down, which combines Holm’s sequentially rejective ordering with Sidak’s per-comparison level. It controls the family-wise error rate while rejecting at least as much as plain Bonferroni, so the correction costs less power than the usual objection to correcting assumes (Holm 1979; Sidak 1967). Sort the p-values ascending and test the -th against
A non-significant result is a failure to detect a difference, never evidence of equivalence. Claiming two checkpoints are the same because a test did not reject requires an equivalence test with a stated margin, which is a different procedure.
Anomaly thresholds under label scarcity
Section titled “Anomaly thresholds under label scarcity”Where labels are too scarce to calibrate directly, triage can still estimate a density over a score and flag low-density runs. A score on the open unit interval breaks a kernel fit directly, since mass leaks past the boundary. Fit on the logit and recover through the Jacobian:
Bandwidth by cross-validated held-out log-likelihood, not a rule of thumb. The Gaussian-kernel CDF over domain peers gives a smoothed percentile:
Typicality is leave-one-out, and the threshold is a percentile of the LOO log-density, itself bootstrapped:
This is triage, not calibration. An anomaly threshold says a run is unusual against its peers; it carries no claim that the run is wrong.
Active labeling
Section titled “Active labeling”Which run to label next combines two leave-one-out signals against : uncertainty, maximal at the decision boundary, and influence, how far removing the run moves the threshold. Each is min-max normalized, then averaged.
Expansion should be gated on inter-annotator agreement clearing a stated floor, because labeling faster than annotators agree buys noise at the price of signal. Multi-rater sheets score Krippendorff’s , which handles unequal raters per item, missing data, and any measurement level, and is the reason a labeling program does not need every annotator to see every item (Hayes and Krippendorff 2007):
What a calibration record has to carry
Section titled “What a calibration record has to carry”A number without these is not checkable, and should not be quoted.
| Field | Why it is load-bearing |
|---|---|
| Sample size and class balance | Fixes what the prior contributes, and what a constant predictor would score |
| Anchor provenance | Human label, or a proxy; a proxy bounds the claim to association with that proxy |
| Annotator count and repeat rate | Without repeats, is unmeasurable and the ceiling above is unknown |
| Fitted or held out | A set the method was tuned on is an implementation diagnostic, never evidence of validity |
| Interval, with its method | Wilson for a proportion, bootstrap for ; a bare point estimate hides the resolution |
| Per-class counts | False accepts and false rejects separately, since accuracy hides their trade |
| Per-class abstention rate | The rates and the summary below are computed on the decided subset, so a record without its coverage cannot be compared with one taken at a different refusal rate |
| A prevalence-free summary | Accuracy moves with the class balance, and so does through its marginals. Sensitivity plus specificity minus one does not, so it is the figure that survives comparison across sets with different priors, on the decided subset (Youden 1950) |
| Baseline row | A constant predictor, showing what the prior alone already buys |
The distinction that most often goes missing is fitted versus held out. A result produced on the set the method was adjusted against measures implementation, not validity, and it belongs in the record labeled as such rather than omitted or promoted.
What blocks calibration
Section titled “What blocks calibration”Usually labels, not model capacity. A dimension checked against a proxy can report a result, but the result is agreement with the proxy rather than validation of the dimension, and no amount of model improvement converts one into the other. A dimension with no human labels at all can compute rates and still has to abstain from any validity claim.
The gate on inter-annotator agreement is the cheapest falsifier for the whole labeling program. Below it, a judge calibrated against either annotator is measuring the annotator rather than the agent.
References
Section titled “References”| Work | Bears on |
|---|---|
| Wilson, E. B. (1927). Probable inference, the law of succession, and statistical inference. Journal of the American Statistical Association 22, 209-212. DOI | The score interval for a proportion |
| McNemar, Q. (1947). Note on the sampling error of the difference between correlated proportions or percentages. Psychometrika 12(2), 153-157. DOI | Paired binary comparison on discordant pairs |
| Youden, W. J. (1950). Index for rating diagnostic tests. Cancer 3, 32-35. DOI | A prevalence-free summary of a binary instrument |
| Smirnov, N. (1948). Table for estimating the goodness of fit of empirical distributions. The Annals of Mathematical Statistics 19, 279-281. DOI | Comparing two distributions rather than two means |
| Bross, I. D. J. (1954). Misclassification in 2 X 2 tables. Biometrics 10, 478-486. DOI | Label noise attenuates a measured association toward the null |
| Cohen, J. (1960). A coefficient of agreement for nominal scales. Educational and Psychological Measurement 20, 37-46. DOI | Chance-corrected agreement |
| Sidak, Z. (1967). Rectangular confidence regions for the means of multivariate normal distributions. Journal of the American Statistical Association 62, 626-633. DOI | The per-comparison level used in the step-down correction |
| Landis, J. R. and Koch, G. G. (1977). The measurement of observer agreement for categorical data. Biometrics 33(1), 159-174. DOI | The interpretation bands, described by their own authors as arbitrary |
| Efron, B. (1979). Bootstrap methods: another look at the jackknife. The Annals of Statistics 7. DOI | Resampling an interval for a statistic with no closed form |
| Dawid, A. P. and Skene, A. M. (1979). Maximum likelihood estimation of observer error-rates using the EM algorithm. Applied Statistics 28(1), 20-28. DOI | Latent-truth aggregation with per-rater confusion matrices |
| Holm, S. (1979). A simple sequentially rejective multiple test procedure. Scandinavian Journal of Statistics 6(2), 65-70. JSTOR | Step-down family-wise error control, uniformly more powerful than Bonferroni |
| Feinstein, A. R. and Cicchetti, D. V. (1990). High agreement but low kappa: I. The problems of two paradoxes. Journal of Clinical Epidemiology 43, 543-549. PubMed | Why a skewed label distribution depresses independently of rater skill |
| Brown, L. D., Cai, T. T. and DasGupta, A. (2001). Interval estimation for a binomial proportion. Statistical Science 16(2), 101-133. Project Euclid | Wald coverage is erratic, not merely weak at the extremes; Wilson is the recommendation |
| Hayes, A. F. and Krippendorff, K. (2007). Answering the call for a standard reliability measure for coding data. Communication Methods and Measures 1, 77-89. DOI | Reliability across unequal raters per item, missing data, and any measurement level |