Release Gate
Between a score and a release decision sits a gate. It checks the score matrix in order, compares releases, and records what evidence supports each dimension. The human labels the gate relies on are specified on LLM Judge.
Decision boundary
Section titled “Decision boundary”The gate checks whether scores can be interpreted, not whether the scored system is good. It evaluates a release, not a live stream.
| Symbol | Meaning |
|---|---|
| score matrix | The table of scored runs: one row per run, one column per axis |
| calibrated | A dimension checked against held-out human labels, with agreement published |
| provisional | A dimension that produces a value with nothing yet supporting its validity |
| label coverage | The share of runs carrying a human anchor for a given axis |
Reading. The matrix is a vector of measurements. No column collapses the axes into a scalar. A regression on an uncalibrated dimension routes to human review, never to an automatic promote.
Gate sequence
Section titled “Gate sequence”Three check families run in order. A later family never rescues an earlier failure. Any hard-check failure fails the combined verdict.
- Data validity. Row count matches the emitted count; every row carries required keys and non-empty task text; identifiers are unique.
- Scoring consistency. The matrix rebuilds deterministically from the same content hash; every value sits in its declared range; the registry aligns with the axis list; degenerate axes are flagged; no column collapses axes into one scalar.
- Interpretability. Each live metric carries signal; label coverage and base rates are reported; every metric claiming calibration carries its evidence; every metric has a documented name and status.
Dimension status
Section titled “Dimension status”A dimension’s lifecycle has two states. Its promotion runs one way. It enters as provisional: it scores, and nothing more is claimed. It leaves provisional only by acquiring a human anchor and publishing agreement against it (calibrated). Promotion never happens by swapping in a stronger model: a stronger model is a new instrument, not new evidence about the old one’s validity. A calibrated dimension that loses its anchor returns to provisional.
Compare releases
Section titled “Compare releases”Every reported rate carries a Wilson score or Clopper-Pearson interval (Brown, Cai, and DasGupta 2001). Its comparator table uses the same population and carries a constant predictor, a random predictor at the base rate, and a cheap structural heuristic, so skew, trace length, and step count cannot read as skill.
Two releases scored on the same runs are compared by McNemar’s exact test over discordant pairs (McNemar 1947). When the run sets differ, a two-proportion comparison is used for binary axes and the loss of pairing is stated. Continuous axes on shared runs are compared through paired per-run differences; report the median difference with a paired-resampling interval. For different run sets, use a two-sample Kolmogorov-Smirnov test plus a median shift (Massey 1951), not a mean difference alone.
With the count of runs only the first release scores correctly and the count only the second release scores correctly, let . The exact two-sided -value is . Concordant runs carry no signal about which release is better.
A family of comparisons uses a Holm-Bonferroni correction (Holm 1979). For valid -values, it controls family-wise error under arbitrary dependence. For ordered -values , compare against
Reading. A result significant before the correction and not after routes to human review. It is not promoted and not silently dropped.
Limit. No shared runs between two releases means insufficient data for a paired claim, never an automatic promote.
Verdict and limits
Section titled “Verdict and limits”| Failure class | Resolution |
|---|---|
| Any hard check fails | The release is blocked |
| Degenerate axis detected | Flagged; never silently dropped |
| Regression on a calibrated dimension | Hard regress; blocked |
| Regression on an uncalibrated dimension | Human review |
| Significant before correction, not after | Human review |
| No shared runs for a comparison | Insufficient data; no paired claim |
The interpretability thresholds are proposed working values, not validated release criteria. They may route a release to human review. They cannot justify an automatic promotion.