Skip to content

What a GUI Agent Is

A GUI agent does its work through the same interface a person would use. It inherits that interface’s constraints and none of the guarantees a private API would give it. Once that is true, the honest unit of measurement is the run, recorded as a trace, judged against the terminal state, not against the agent’s own report of success. Scoring dimensions, rubrics, and judges live in Evaluation. This page only names the objects those instruments read.

The agent has no private channel into the application. It sees what is on the screen, chooses an action a person could take (tap, type, scroll, press a key), and waits for the interface to respond. Rendering, focus, timing, and layout are part of the problem.

Observe, decide, act, check loop with partial observability environment state visible slice most state stays hidden observe agent decide step cycle act check did the act land? return to environment
Figure 1.The loop. Only a slice of environment state is visible. Observation enters the agent; action and check return to the environment. Animated edges are emphasis only; the diagram reads without motion.
Soft illustration of a large interface with a smaller sharp viewport showing only a slice of controls.
Figure 2.Illustrative plate. Most of the interface stays outside the agent's sharp window. Atmosphere only; not a measurement.

A recurring failure in this class of system is a confident success message over a screen that still shows the task unfinished. Score what the screen shows at the end, not a moment mid-run that briefly looked right. If success cannot be pointed at in a recorded frame, it has not been shown.

Warm paper illustration of a phone showing an incomplete screen beside a Done status chip.
Figure 3.Illustrative plate. The screen is the evidence; a Done claim beside an unfinished screen is not. Atmosphere only; not a measurement.

Two runs can reach the same end state in seven steps or in fifty. One of them may have fired a purchase, a send, a delete, or a booking without confirmation. The end screen alone cannot tell those runs apart; the ordered steps can.

Reading. After the fact, the trace is the only object available to audit or score. The ordered steps are the route. Collapse the run to a single pass or fail and that route averages away.

ObjectMeaning
RunOne attempt at one task, from the opening screen until the agent stops
TraceThe record of a run
StepOne observe, decide, act cycle
Terminal stateEnd screen; evidence for completion
Consequential actionStep that needs authorization before it fires

Whether one run succeeded is a Bernoulli outcome. A set of nn independent runs with kk successes yields a binomial proportion p^=k/n\hat{p} = k/n. A proportion without its nn and an interval is not yet a result (Brown, Cai, and DasGupta 2001).

Per-step success and per-run success are different quantities. Under an independence assumption, if each step succeeds with probability pp, then a run of nn steps succeeds with probability

P(run succeeds)=pn.P(\text{run succeeds}) = p^{n}.

Reading. A high per-step number and a low per-run number can describe the same agent, because length compounds. Reporting only step accuracy flatters systems that fail on long tasks.

Limit. Steps in a live interface are not independent: an error changes the state the next step observes. Treat pnp^{n} as an intuition pump for why length matters, not as a deployed rate (Ross and Bagnell 2010).

Per-step success compounds into lower run success as length grows illustrative: p = 0.95 per step, independent 100% 0% 77% 5 steps 60% 10 steps 36% 20 steps 13% 40 steps run success under independence (intuition only)
Figure 4.Illustrative compounding at p = 0.95 under independence. Bar heights are computed in the figure from that p and the step counts shown. Not a measured rate; independence is the wrong assumption in practice.

If one number cannot carry the run, the output has to be several, and the unit of report is the run with its trace. How those numbers are named, scored, and gated is Evaluation’s job. What the agent can see and do, and how a long run holds together, are the next two pages in this section.

WorkCitation
Brown, L. D., Cai, T. T., and DasGupta, A. (2001). Interval Estimation for a Binomial Proportion. Statistical Science 16(2), 101-133.
Ross, S. and Bagnell, J. A. (2010). Efficient Reductions for Imitation Learning. Proceedings of the Thirteenth International Conference on Artificial Intelligence and Statistics (AISTATS), PMLR 9:661-668.