LLM Evaluation: An Introduction
An LLM does not produce a verdict of itself; a measurement instrument does. This page states what that instrument is, what it must score, and why collapsing to a single number hides the result. Scoring dimensions, rubrics, and judges live here in Evaluation. The run and the trace, as runtime objects, are named in What a GUI Agent Is.
Three objects
Section titled “Three objects”An evaluation scores one of three different things. Conflating them is the first way a measurement goes wrong.
- Model outputs. Whether a response is good on its own terms, judged against a rubric for the task it was given.
- Agent routes. Whether the route to the answer was sound. Two agents can reach the same end screen by different paths, and the path predicts whether the next task holds up. The record of that route is the trace.
- Judges. Whether the instrument itself is trustworthy. The first two are what most teams mean by evaluation; the third is the one they skip, and it is the one that decides whether the other two mean anything.
Reading. A strong model answer, a sound agent route, and a calibrated judge are three different claims. Publish which one you measured.
Why a single number fails
Section titled “Why a single number fails”A single pass or fail label hides the run. Two runs can both reach the goal, one in seven steps and one in fifty. One confirms before deleting, one does not. One acts on the screen in front of it, one acts on a stale guess. Collapse each to pass or fail and the differences average away.
A recurring failure in this class of system is a confident completion claim the human rejects when looking at the end screen. A pass-biased single label hides exactly that pattern.
The output of a good evaluation is therefore a vector over named dimensions, never a scalar. A run that succeeded but wandered shows as low efficiency instead of disappearing into a pass. A run that succeeded and was unsafe shows as failed safety. Each dimension is one answer to one question, and the questions stay separate.
Limit. Naming four dimensions does not prove they are independent or equally important. Weights, if used at all, are fixed per task family and versioned with the rubric. They do not erase a failed safety dimension.
Verdict space
Section titled “Verdict space”Every dimension returns a value from a small, fixed menu. The menu has two decisions (pass, fail) and four non-decisions (abstain, not applicable, missing evidence, not implemented). The four non-decisions are never coerced into pass or fail.
Score and coverage
Section titled “Score and coverage”For a vector of dimension scores, the item report is the vector together with its coverage.
Reading. For ordinal or continuous scores, the same shape applies: a score plus a coverage, where coverage is the fraction of dimensions the instrument actually decided. An instrument that abstains on a third of its dimensions is reporting a different number than one that decided everything.
A high score at low coverage is a weaker claim than the same score at full coverage. A change to how dimensions are weighted, if weights are used, is a change to the standard and is versioned alongside the rubric and the model. The vector and its coverage, not a single blended number, are the reportable unit (Brown, Cai, and DasGupta 2001).
Limit. Coverage is not agreement. Coverage says how often the instrument decided; agreement says how often those decisions match a gold set. Both appear on the scorecard; neither substitutes for the other (Zheng 2023).
The pipeline that turns a recorded run into that reportable unit is the next page. Rubrics, taxonomies, and the judge itself follow after.