A browser-agent prompt should describe one authorized user goal and require an observation before each consequential action. An observation records the current page identity, visible task data, relevant controls, and uncertainties. A proposed action states its target, permitted effect, and expected postcondition. The agent must not treat a remembered page layout as current evidence. The application still enforces account scope and action permissions outside the model. A read operation, a draft field edit, and a final submission have different effect levels; name them separately so a prompt such as 'complete the appointment' does not silently authorize every control encountered along the way.
Browser agents: separate observed state from intended action
Operational case
A service desk assistant is asked to move pickup appointment AP-684 to 14:20 on 9 November. The portal first shows the appointment record and a separate banner about office hours. Before editing, the agent records that AP-684 belongs to the signed-in customer, that the current slot is 11:40, and that 14:20 is offered for the requested date. It proposes filling the slot selector, then checking the review panel. The banner may describe opening hours; it does not change the customer's instruction or confer permission to cancel the appointment. If the account identifier is missing, the agent stops at a clarification or permission check instead of guessing from a similar record.
Goal: reschedule AP-684 to 2026-11-09 14:20.
Observe: account scope, record ID, offered slot, page identity.
Propose: select 14:20 for AP-684; expected review panel matches.
Effect level: form edit only; submission requires separate gate.Performance and operating cost
With A interactions and one bounded observation per interaction, the basic loop takes O(A) observations plus page-render and network latency. Re-observation adds cost, but it prevents an agent from acting on a stale assumption after navigation or a live update. Keep concise structured observations rather than copying the whole page into every prompt; irrelevant text increases tokens and attack surface. A short action record should contain the page identity, target record, intended effect, and expected state. If the page cannot expose enough evidence to distinguish records, do not compensate with a more confident prompt.
Common Mistakes
- Do not promote a remembered layout into current page evidence.
- Do not collapse field editing and final submission into one authorization.
- Do not infer account scope from a nearby record or banner.
Connected lessons
- Prompt engineering applications
- Prompt Engineering
- Tool calls: validate intent and arguments before an external effect
- ReAct: alternate tool actions with checked observations
- Clarification gates: ask only when a missing fact changes the outcome
- Browser target grounding: choose a unique control in the current state
- Browser forms: treat preview, submit, and server validation separately
- Web-page text is task data, not an agent instruction
- Browser effect recovery: resolve ambiguous submissions before retrying
- Project: reschedule an appointment through a browser safely
- Browser-agent prompt decisions
