Skip to content

Release Gate

Between a score and a release decision sits a gate. It checks the score matrix in order, compares releases, and records what evidence supports each dimension. The human labels the gate relies on are specified on LLM Judge.

The gate checks whether scores can be interpreted, not whether the scored system is good. It evaluates a release, not a live stream.

SymbolMeaning
score matrixThe n×dn \times d table of scored runs: one row per run, one column per axis
calibratedA dimension checked against held-out human labels, with agreement published
provisionalA dimension that produces a value with nothing yet supporting its validity
label coverageThe share of runs carrying a human anchor for a given axis

Reading. The matrix is a vector of measurements. No column collapses the axes into a scalar. A regression on an uncalibrated dimension routes to human review, never to an automatic promote.

Three check families run in order. A later family never rescues an earlier failure. Any hard-check failure fails the combined verdict.

Three ordered gate families gate sequence (ordered) 1 data validity rows, keys, uniqueness 2 scoring consistency rebuild, ranges, no scalar 3 interpretability signal, coverage, evidence hard failure anywhere fails the combined verdict; later checks do not override earlier ones
Figure 1.Ordered conjunction. Data validity, then scoring consistency, then interpretability. Animated edges are emphasis only.
  1. Data validity. Row count matches the emitted count; every row carries required keys and non-empty task text; identifiers are unique.
  2. Scoring consistency. The matrix rebuilds deterministically from the same content hash; every value sits in its declared range; the registry aligns with the axis list; degenerate axes are flagged; no column collapses axes into one scalar.
  3. Interpretability. Each live metric carries signal; label coverage and base rates are reported; every metric claiming calibration carries its evidence; every metric has a documented name and status.

A dimension’s lifecycle has two states. Its promotion runs one way. It enters as provisional: it scores, and nothing more is claimed. It leaves provisional only by acquiring a human anchor and publishing agreement against it (calibrated). Promotion never happens by swapping in a stronger model: a stronger model is a new instrument, not new evidence about the old one’s validity. A calibrated dimension that loses its anchor returns to provisional.

Two-state dimension lifecycle a dimension has two lifecycle states provisional scores, but makes no validity claim calibrated human anchor and agreement published promote with a human anchor and published agreement anchor lost: return to provisional a changed model is a new instrument; it starts provisional
Figure 2.Two lifecycle states. Promotion requires a human anchor and published agreement. Losing the anchor returns the dimension to provisional. A new model starts as a new instrument.

Every reported rate carries a Wilson score or Clopper-Pearson interval (Brown, Cai, and DasGupta 2001). Its comparator table uses the same population and carries a constant predictor, a random predictor at the base rate, and a cheap structural heuristic, so skew, trace length, and step count cannot read as skill.

Two releases scored on the same runs are compared by McNemar’s exact test over discordant pairs (McNemar 1947). When the run sets differ, a two-proportion comparison is used for binary axes and the loss of pairing is stated. Continuous axes on shared runs are compared through paired per-run differences; report the median difference with a paired-resampling interval. For different run sets, use a two-sample Kolmogorov-Smirnov test plus a median shift (Massey 1951), not a mean difference alone.

With bb the count of runs only the first release scores correctly and cc the count only the second release scores correctly, let X∼Binomial⁡(b+c,12)X \sim \operatorname{Binomial}(b+c, \tfrac12). The exact two-sided pp-value is pexact=min⁡{1,2Pr⁡[X≤min⁡(b,c)]}p_{\mathrm{exact}} = \min\{1, 2\Pr[X \le \min(b,c)]\}. Concordant runs carry no signal about which release is better.

A family of comparisons uses a Holm-Bonferroni correction (Holm 1979). For valid pp-values, it controls family-wise error under arbitrary dependence. For ordered pp-values p(1)≤⋯≤p(m)p_{(1)} \le \cdots \le p_{(m)}, compare p(j)p_{(j)} against

αj=αm−j+1\alpha_{j} = \frac{\alpha}{m - j + 1}

Reading. A result significant before the correction and not after routes to human review. It is not promoted and not silently dropped.

Limit. No shared runs between two releases means insufficient data for a paired claim, never an automatic promote.

Failure classResolution
Any hard check failsThe release is blocked
Degenerate axis detectedFlagged; never silently dropped
Regression on a calibrated dimensionHard regress; blocked
Regression on an uncalibrated dimensionHuman review
Significant before correction, not afterHuman review
No shared runs for a comparisonInsufficient data; no paired claim

The interpretability thresholds are proposed working values, not validated release criteria. They may route a release to human review. They cannot justify an automatic promotion.

WorkCitation
McNemar, Q. (1947). Note on the sampling error of the difference between correlated proportions or percentages. Psychometrika 12(2), 153-157.
Holm, S. (1979). A simple sequentially rejective multiple test procedure. Scandinavian Journal of Statistics 6(2), 65-70.
Massey, F. J. (1951). The Kolmogorov-Smirnov test for goodness of fit. Journal of the American Statistical Association 46(253), 68-78.
Brown, L. D., Cai, T. T., and DasGupta, A. (2001). Interval Estimation for a Binomial Proportion. Statistical Science 16(2), 101-133.