Failures
When a scored dimension fails, the judge writes one failure record. The record carries the run’s identity and provenance, then describes what failed, how, who contributed, when it took shape, and what evidence supports each assertion. Guide diagnostics are kept separate: they route repair but do not change the score vector. This page states the taxonomy, the assignment rules, the falsifier a new type must pass, and how the record points at a repair.
Taxonomy
Section titled “Taxonomy”The taxonomy has two levels: family and type. The family is derived from the type, never stored beside it, so nothing can drift. Each type belongs to one failed scored dimension. Cross-links are analytical; they do not add a score.
| Family | Types (what failed) | Scored dimension |
|---|---|---|
| Completion | missing, wrong, extra | completion |
| Safety | unsafe outcome | safety |
| Efficiency | wasteful | efficiency |
| Grounding | ungrounded step | grounding |
Guide diagnostics route repair after a scored failure is recorded. They are not failure types and do not add a fifth scored dimension.
| Diagnostic lane | Findings | Role |
|---|---|---|
| Guide | retrieval miss, fidelity loss, adherence break, a scored failure remains after adherence | Repair routing only; not a score |
Reading. Naming the type answers what failed. A Guide diagnostic answers which repair to inspect. Attribution answers whose contribution to investigate. Mixing them produces a label nobody can act on.
Obligation alignment
Section titled “Obligation alignment”The three completion types are not guesses; they fall out of obligation alignment. Decompose the task into the terminal facts it requires, then compare each to the final screen.
Reading. Safety and efficiency are judged over the same run on their own dimensions, not on this alignment. A run can be complete and unsafe.
Limit. Alignment assumes the task’s terminal obligations can be stated. Vague goals yield vague types; fix the task statement before trusting the completion label.
Layers of a failure record
Section titled “Layers of a failure record”When a scored dimension fails, the judge writes one record. Identity and provenance sit beside six descriptive layers. A guide diagnostic, when applicable, routes repair without altering the score vector.
| Layer | Question |
|---|---|
| Identity | Which record is this? (append-only; a rerun is a new record) |
| Provenance | Which inputs could change the output? |
| Dimension | Which scored property failed? |
| Type | What is the observable form? |
| Mechanism | How did it happen? (optional, model-judged) |
| Attribution | Who contributed? (multi-label) |
| Trajectory | When in the run? (deterministic pattern) |
| Evidence | What supports each assertion? |
Record rules that keep the vocabulary honest:
- The family is derived from the type, never stored beside it.
- Evidence attaches per assertion, not per record.
- Attribution is a set; it is not a failure type.
- A trajectory pattern answers when, not what.
- Mechanism is optional; unknown is recorded as unknown, not guessed.
- A run below the evidence floor (terminal state unreadable) yields missing evidence, which is not a failure type.
- Guide diagnostics route repair; they never replace a failed dimension or a type.
Once the four rubric dimensions are scored, assignment is mechanical: completion routes through obligation alignment; a failed safety, efficiency, or grounding dimension carries one type; and trajectory and attribution are added as separate layers. A guide diagnostic may route repair, but it cannot carry a score.
Trajectory categories include loop, premature finish, timeout, handoff, reversal, and residual. Residual is the not-yet-diagnosed remainder, kept visible so coverage of the rule is visible. Attribution is multi-label: agent, guide, environment, capture, scorer, or infrastructure can share a failure.
Falsifier
Section titled “Falsifier”A candidate becomes a named type only if it passes the discovery falsifier. Clustering for discovery runs only on what the failure signature does not already name. A two-dimensional embedding may draw the scatter; it must not do the grouping.
A candidate type must (a) reappear on a disjoint run set, (b) survive a different device and application mix, (c) survive removal of trajectory-length features, (d) survive removal of family-label text features, and (e) be assigned independently by two blinded annotators, using the written definition alone, in a pilot that clears the agreement bar on Rubric. A candidate that dissolves under any clause is a feature-space artifact, not a type (Campello et al. 2015; McInnes et al. 2018).
| Failure class | Resolution |
|---|---|
| Run below the evidence floor | Record missing evidence; no type assigned |
| Run fits no trajectory rule | Trajectory is residual |
| Cluster candidate fails a falsifier clause | Discarded as a feature-space artifact |
| Mechanism undecidable | Recorded as unknown; the type stands without it |
Limit. How well the types cover real failures, and how separable they are under independent labeling, is unmeasured until a labeled coverage study runs. Secondary links between types are analytical until that study lands.
Position and repair
Section titled “Position and repair”The score vector and repair routing do different jobs. Completion, safety, efficiency, and grounding are the scored dimensions. Guide diagnostics locate a repair review for a scored failure. They do not add a fifth dimension or prove a cause.
Apply the checks in order. First ask whether terminal evidence clears the evidence floor. If no guide was served, check grounding before routing a missing intended guide to retrieval review. For a served guide, ask whether it was faithful and followed. A scored failure that remains after adherence sends the guide to authoring review. If no diagnostic applies, retain the scored type and investigate attribution.
Reading. An unreadable final frame is missing evidence, not an agent failure. A scored failure that remains after adherence sends the guide to authoring review; it does not establish that the guide caused the failure.
Limit. Repair routing localizes the first inspection. Attribution remains multi-label because agent, guide, environment, capture, scorer, and infrastructure can still contribute together.
Human anchors for the four scored dimensions are produced under LLM Judge.