What the Agent Can See and Do
Handing the agent one object called the screen state is the natural design and the expensive mistake. Split the view instead: what was literally read off the frame, and what those readings were taken to mean. Everything downstream should consume only the second layer, so the first stays available as evidence. Which properties get scored belongs to Evaluation. This page states the runtime objects Evaluation later attributes to.
Two layers
Section titled “Two layers”Once the layers share a namespace, a perception miss and a planning miss produce the same kind of record, and the fix gets routed to the wrong place. The split has to be checked at runtime, not only written as a convention.
Reading. Downstream policy and scoring should read the inferred state. The raw layer stays as evidence. Mixing the two namespaces promotes a misread into fact.
Attribution in the failure record (agent, environment, capture, and related labels) is the Evaluation-side consumer of this split.
Actions
Section titled “Actions”Reading. Recording the allowed set beside the step turns an unauthorized action into a checkable event. Taps, typing, drags, launches, and navigation are gradable against a screen; waiting, finishing, and handing control back are endings or pauses, not grounded moves.
Perception audit
Section titled “Perception audit”The accessibility tree the operating system renders is an independent view of the same screen. It is not inferred from pixels, so disagreement between the two views on the same control is a perception fault, not an agent-policy failure.
Re-reading the same frame and measuring agreement among readings measures stability, not correctness: a reader that is consistently wrong agrees with itself. Known-answer frames catch systematic error. Where an app suppresses accessibility reporting, record that absence rather than treating it as benign.
Platform shows through
Section titled “Platform shows through”Verbs are close to universal across desktop, mobile, and web. What varies is how a target is named, and quiet harness details that break comparability: coordinate normalisation, screen size, wait policy, retry budget. The adapter preserves the action set and the meaning of a target, not the look of the pixels. Record the adapter or platform version beside the trace.
| Object | Meaning |
|---|---|
| Raw observation | Literal read of one frame |
| Inferred state | What those readings imply |
| Action | Verb, target, and action class |
| Allowed set | Action classes permitted at this step |
| Target reference | How the action names its target |
The next page asks what happens when a task outlasts what the agent can see or hold at once.