Skip to content

Planning Across Long Tasks

An agent that takes ten correct steps in a row is not the same thing as an agent that finishes a long task. The gap is where most real failures sit, and the useful place to start is how the run ended. Names and counts for those failures belong to Evaluation. This page states the runtime objects: endings, drift, what is carried forward, and honest stops.

Short tasks often fail on grounding. Long tasks fail on drift, loops, and beliefs formed early and never rechecked. A useful sketch puts a cost on each transition between consecutive inferred states, high when nothing on the earlier screen licenses the later one. A loop is periodic structure in the per-step series; the average destroys exactly that structure.

Transition cost along a run: jump then loop illustrative route shape steady jump loop per-step transition cost (shape carries what an average hides)
Figure 1.Illustrative route shape. Early steps stay low, one tall bar marks an unlicensed jump, then a repeating pattern shows a loop. Bar heights are geometry for teaching, not measured costs. Animation is emphasis only.

Reading. The ending belongs in the trace as its own field. Without it, a failure and a correct decision to stop look the same downstream. Residual (not yet diagnosed) keeps the coverage of the rule visible. Handing back is a first-class ending, not a failure.

Four defensible endings: the goal is visible on screen, the step budget is spent, control is handed back, or the next step is refused.

Reading. A plan is a short list of subgoals, not a script. Three triggers carry most of the weight: the expected screen did not appear, the same state repeated, and the step budget is nearly spent.

A retrieved procedure fails in four ways: the wrong one was fetched, key steps were lost when it was condensed, the run did not follow it, or following it did not help. Evaluation may score those separately; this page only names the runtime points. A confidently stale belief is worse than none, so carried state needs a written-at marker and a drop rule. A longer context window does not remove the problem (Liu et al. 2023).

ObjectMeaning
Step budgetCap on steps; timeout when hit unfinished
Trajectory patternRule-assigned temporal shape (Evaluation when-layer)
Working setState carried step to step
Retrieved procedureGuidance fetched once for the run
HandoffReturn of control to the person
WorkCitation
Liu, N. F., et al. (2023). Lost in the Middle: How Language Models Use Long Contexts. arXiv:2307.03172.