Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Browser agents: separate observed state from intended action

Last updated: 5 Oct 202611 min read
tutorial
AdvancedBy AITrove Editorial

A browser-agent prompt should describe one authorized user goal and require an observation before each consequential action. An observation records the current page identity, visible task data, relevant controls, and uncertainties. A proposed action states its target, permitted effect, and expected postcondition. The agent must not treat a remembered page layout as current evidence. The application still enforces account scope and action permissions outside the model. A read operation, a draft field edit, and a final submission have different effect levels; name them separately so a prompt such as 'complete the appointment' does not silently authorize every control encountered along the way.

Operational case

A service desk assistant is asked to move pickup appointment AP-684 to 14:20 on 9 November. The portal first shows the appointment record and a separate banner about office hours. Before editing, the agent records that AP-684 belongs to the signed-in customer, that the current slot is 11:40, and that 14:20 is offered for the requested date. It proposes filling the slot selector, then checking the review panel. The banner may describe opening hours; it does not change the customer's instruction or confer permission to cancel the appointment. If the account identifier is missing, the agent stops at a clarification or permission check instead of guessing from a similar record.

Output
Goal: reschedule AP-684 to 2026-11-09 14:20.
Observe: account scope, record ID, offered slot, page identity.
Propose: select 14:20 for AP-684; expected review panel matches.
Effect level: form edit only; submission requires separate gate.

Performance and operating cost

With A interactions and one bounded observation per interaction, the basic loop takes O(A) observations plus page-render and network latency. Re-observation adds cost, but it prevents an agent from acting on a stale assumption after navigation or a live update. Keep concise structured observations rather than copying the whole page into every prompt; irrelevant text increases tokens and attack surface. A short action record should contain the page identity, target record, intended effect, and expected state. If the page cannot expose enough evidence to distinguish records, do not compensate with a more confident prompt.

Common Mistakes

  • Do not promote a remembered layout into current page evidence.
  • Do not collapse field editing and final submission into one authorization.
  • Do not infer account scope from a nearby record or banner.

Connected lessons

prompt engineering
browser agents
Storage details