What a GUI Agent Is
A GUI agent does its work through the same interface a person would use. It inherits that interface’s constraints and none of the guarantees a private API would give it. Once that is true, the honest unit of measurement is the run, recorded as a trace, judged against the terminal state, not against the agent’s own report of success. Scoring dimensions, rubrics, and judges live in Evaluation. This page only names the objects those instruments read.
Interface constraint
Section titled “Interface constraint”The agent has no private channel into the application. It sees what is on the screen, chooses an action a person could take (tap, type, scroll, press a key), and waits for the interface to respond. Rendering, focus, timing, and layout are part of the problem.
A recurring failure in this class of system is a confident success message over a screen that still shows the task unfinished. Score what the screen shows at the end, not a moment mid-run that briefly looked right. If success cannot be pointed at in a recorded frame, it has not been shown.
Two runs can reach the same end state in seven steps or in fifty. One of them may have fired a purchase, a send, a delete, or a booking without confirmation. The end screen alone cannot tell those runs apart; the ordered steps can.
Reading. After the fact, the trace is the only object available to audit or score. The ordered steps are the route. Collapse the run to a single pass or fail and that route averages away.
| Object | Meaning |
|---|---|
| Run | One attempt at one task, from the opening screen until the agent stops |
| Trace | The record of a run |
| Step | One observe, decide, act cycle |
| Terminal state | End screen; evidence for completion |
| Consequential action | Step that needs authorization before it fires |
What measurement inherits
Section titled “What measurement inherits”Whether one run succeeded is a Bernoulli outcome. A set of independent runs with successes yields a binomial proportion . A proportion without its and an interval is not yet a result (Brown, Cai, and DasGupta 2001).
Per-step success and per-run success are different quantities. Under an independence assumption, if each step succeeds with probability , then a run of steps succeeds with probability
Reading. A high per-step number and a low per-run number can describe the same agent, because length compounds. Reporting only step accuracy flatters systems that fail on long tasks.
Limit. Steps in a live interface are not independent: an error changes the state the next step observes. Treat as an intuition pump for why length matters, not as a deployed rate (Ross and Bagnell 2010).
If one number cannot carry the run, the output has to be several, and the unit of report is the run with its trace. How those numbers are named, scored, and gated is Evaluation’s job. What the agent can see and do, and how a long run holds together, are the next two pages in this section.