Skip to content

Failures

When a scored dimension fails, the judge writes one failure record. The record carries the run’s identity and provenance, then describes what failed, how, who contributed, when it took shape, and what evidence supports each assertion. Guide diagnostics are kept separate: they route repair but do not change the score vector. This page states the taxonomy, the assignment rules, the falsifier a new type must pass, and how the record points at a repair.

The taxonomy has two levels: family and type. The family is derived from the type, never stored beside it, so nothing can drift. Each type belongs to one failed scored dimension. Cross-links are analytical; they do not add a score.

FamilyTypes (what failed)Scored dimension
Completionmissing, wrong, extracompletion
Safetyunsafe outcomesafety
Efficiencywastefulefficiency
Groundingungrounded stepgrounding

Guide diagnostics route repair after a scored failure is recorded. They are not failure types and do not add a fifth scored dimension.

Diagnostic laneFindingsRole
Guideretrieval miss, fidelity loss, adherence break, a scored failure remains after adherenceRepair routing only; not a score

Reading. Naming the type answers what failed. A Guide diagnostic answers which repair to inspect. Attribution answers whose contribution to investigate. Mixing them produces a label nobody can act on.

The three completion types are not guesses; they fall out of obligation alignment. Decompose the task into the terminal facts it requires, then compare each to the final screen.

Obligation alignment yields missing, wrong, or extra task obligations slot: recipient slot: amount slot: date align end screen facts recipient: Ada amount: wrong extra: tip added missing (empty) wrong (value) extra (no slot) completion types fall out of alignment; safety and efficiency are scored separately
Figure 1.Obligation alignment. Empty slot: missing. Filled slot with the wrong value: wrong. Agent-caused fact with no slot: extra. Illustrative; not a measured run.

Reading. Safety and efficiency are judged over the same run on their own dimensions, not on this alignment. A run can be complete and unsafe.

Limit. Alignment assumes the task’s terminal obligations can be stated. Vague goals yield vague types; fix the task statement before trusting the completion label.

When a scored dimension fails, the judge writes one record. Identity and provenance sit beside six descriptive layers. A guide diagnostic, when applicable, routes repair without altering the score vector.

Layers of one failure record one failure, separate layers identity append-only id provenance what can change the output dimension which property failed type observable form mechanism how (optional) attribution who contributed trajectory when in the run evidence what supports each claim
Figure 2.Identity and provenance, then dimension, type, mechanism, attribution, trajectory, and evidence. Each layer answers a different question.
LayerQuestion
IdentityWhich record is this? (append-only; a rerun is a new record)
ProvenanceWhich inputs could change the output?
DimensionWhich scored property failed?
TypeWhat is the observable form?
MechanismHow did it happen? (optional, model-judged)
AttributionWho contributed? (multi-label)
TrajectoryWhen in the run? (deterministic pattern)
EvidenceWhat supports each assertion?

Record rules that keep the vocabulary honest:

  1. The family is derived from the type, never stored beside it.
  2. Evidence attaches per assertion, not per record.
  3. Attribution is a set; it is not a failure type.
  4. A trajectory pattern answers when, not what.
  5. Mechanism is optional; unknown is recorded as unknown, not guessed.
  6. A run below the evidence floor (terminal state unreadable) yields missing evidence, which is not a failure type.
  7. Guide diagnostics route repair; they never replace a failed dimension or a type.

Once the four rubric dimensions are scored, assignment is mechanical: completion routes through obligation alignment; a failed safety, efficiency, or grounding dimension carries one type; and trajectory and attribution are added as separate layers. A guide diagnostic may route repair, but it cannot carry a score.

Trajectory categories include loop, premature finish, timeout, handoff, reversal, and residual. Residual is the not-yet-diagnosed remainder, kept visible so coverage of the rule is visible. Attribution is multi-label: agent, guide, environment, capture, scorer, or infrastructure can share a failure.

A candidate becomes a named type only if it passes the discovery falsifier. Clustering for discovery runs only on what the failure signature does not already name. A two-dimensional embedding may draw the scatter; it must not do the grouping.

A candidate type must (a) reappear on a disjoint run set, (b) survive a different device and application mix, (c) survive removal of trajectory-length features, (d) survive removal of family-label text features, and (e) be assigned independently by two blinded annotators, using the written definition alone, in a pilot that clears the agreement bar on Rubric. A candidate that dissolves under any clause is a feature-space artifact, not a type (Campello et al. 2015; McInnes et al. 2018).

Failure classResolution
Run below the evidence floorRecord missing evidence; no type assigned
Run fits no trajectory ruleTrajectory is residual
Cluster candidate fails a falsifier clauseDiscarded as a feature-space artifact
Mechanism undecidableRecorded as unknown; the type stands without it

Limit. How well the types cover real failures, and how separable they are under independent labeling, is unmeasured until a labeled coverage study runs. Secondary links between types are analytical until that study lands.

The score vector and repair routing do different jobs. Completion, safety, efficiency, and grounding are the scored dimensions. Guide diagnostics locate a repair review for a scored failure. They do not add a fifth dimension or prove a cause.

First-match repair routing first applicable condition selects the repair lane trace condition diagnostic inspect first 1 terminal evidence unreadable missing evidence capture review 2 no guide served and grounding fails grounding failure training review 3 intended guide not served retrieval miss guide serving 4 served guide is not faithful fidelity loss guide serving 5 agent did not follow a faithful guide adherence break agent policy 6 scored failure remains after adherence failure remains guide authoring review no match: retain the scored type and investigate attribution routing selects an inspection order; it does not establish causation
Figure 3.Illustrative repair routing. The first applicable condition selects one inspection lane. It does not prove cause or change the score.

Apply the checks in order. First ask whether terminal evidence clears the evidence floor. If no guide was served, check grounding before routing a missing intended guide to retrieval review. For a served guide, ask whether it was faithful and followed. A scored failure that remains after adherence sends the guide to authoring review. If no diagnostic applies, retain the scored type and investigate attribution.

Reading. An unreadable final frame is missing evidence, not an agent failure. A scored failure that remains after adherence sends the guide to authoring review; it does not establish that the guide caused the failure.

Limit. Repair routing localizes the first inspection. Attribution remains multi-label because agent, guide, environment, capture, scorer, and infrastructure can still contribute together.

Human anchors for the four scored dimensions are produced under LLM Judge.

WorkCitation
Campello et al. (2015). Hierarchical density estimates for data clustering, visualization, and outlier detection. ACM Transactions on Knowledge Discovery from Data 10(1).
McInnes et al. (2018). UMAP: Uniform Manifold Approximation and Projection. Journal of Open Source Software 3(29), 861.