Skip to content

What the Agent Can See and Do

Handing the agent one object called the screen state is the natural design and the expensive mistake. Split the view instead: what was literally read off the frame, and what those readings were taken to mean. Everything downstream should consume only the second layer, so the first stays available as evidence. Which properties get scored belongs to Evaluation. This page states the runtime objects Evaluation later attributes to.

Once the layers share a namespace, a perception miss and a planning miss produce the same kind of record, and the fix gets routed to the wrong place. The split has to be checked at runtime, not only written as a convention.

Two-layer split of what the agent sees frame one capture raw observation text, boxes, controls inferred state app, screen, focus not consumed downstream decide / act reads inferred only raw stays as evidence for attribution
Figure 1.One frame forks into raw observation and inferred state. Downstream decide and act read only the inferred layer. The raw layer stays as evidence. Animated edges are emphasis only.

Reading. Downstream policy and scoring should read the inferred state. The raw layer stays as evidence. Mixing the two namespaces promotes a misread into fact.

Attribution in the failure record (agent, environment, capture, and related labels) is the Evaluation-side consumer of this split.

Reading. Recording the allowed set beside the step turns an unauthorized action into a checkable event. Taps, typing, drags, launches, and navigation are gradable against a screen; waiting, finishing, and handing control back are endings or pauses, not grounded moves.

The accessibility tree the operating system renders is an independent view of the same screen. It is not inferred from pixels, so disagreement between the two views on the same control is a perception fault, not an agent-policy failure.

Re-reading the same frame and measuring agreement among readings measures stability, not correctness: a reader that is consistently wrong agrees with itself. Known-answer frames catch systematic error. Where an app suppresses accessibility reporting, record that absence rather than treating it as benign.

Verbs are close to universal across desktop, mobile, and web. What varies is how a target is named, and quiet harness details that break comparability: coordinate normalisation, screen size, wait policy, retry budget. The adapter preserves the action set and the meaning of a target, not the look of the pixels. Record the adapter or platform version beside the trace.

ObjectMeaning
Raw observationLiteral read of one frame
Inferred stateWhat those readings imply
ActionVerb, target, and action class
Allowed setAction classes permitted at this step
Target referenceHow the action names its target

The next page asks what happens when a task outlasts what the agent can see or hold at once.