Planning Across Long Tasks
An agent that takes ten correct steps in a row is not the same thing as an agent that finishes a long task. The gap is where most real failures sit, and the useful place to start is how the run ended. Names and counts for those failures belong to Evaluation. This page states the runtime objects: endings, drift, what is carried forward, and honest stops.
How runs end
Section titled “How runs end”Short tasks often fail on grounding. Long tasks fail on drift, loops, and beliefs formed early and never rechecked. A useful sketch puts a cost on each transition between consecutive inferred states, high when nothing on the earlier screen licenses the later one. A loop is periodic structure in the per-step series; the average destroys exactly that structure.
Reading. The ending belongs in the trace as its own field. Without it, a failure and a correct decision to stop look the same downstream. Residual (not yet diagnosed) keeps the coverage of the rule visible. Handing back is a first-class ending, not a failure.
Four defensible endings: the goal is visible on screen, the step budget is spent, control is handed back, or the next step is refused.
What is carried forward
Section titled “What is carried forward”Reading. A plan is a short list of subgoals, not a script. Three triggers carry most of the weight: the expected screen did not appear, the same state repeated, and the step budget is nearly spent.
A retrieved procedure fails in four ways: the wrong one was fetched, key steps were lost when it was condensed, the run did not follow it, or following it did not help. Evaluation may score those separately; this page only names the runtime points. A confidently stale belief is worse than none, so carried state needs a written-at marker and a drop rule. A longer context window does not remove the problem (Liu et al. 2023).
| Object | Meaning |
|---|---|
| Step budget | Cap on steps; timeout when hit unfinished |
| Trajectory pattern | Rule-assigned temporal shape (Evaluation when-layer) |
| Working set | State carried step to step |
| Retrieved procedure | Guidance fetched once for the run |
| Handoff | Return of control to the person |
References
Section titled “References”| Work | Citation |
|---|---|
| Liu, N. F., et al. (2023). Lost in the Middle: How Language Models Use Long Contexts. arXiv:2307.03172. |