Review a fictional Parcel Desk assistant with cross-session memory. North Depot manager U-47 asks for metric units in all future dispatch summaries. A later request uses miles for shipment S-47 only. U-47 also changes a durable delivery-notice preference from email to text in a later explicit statement, then asks to forget the old item. South Depot manager U-84 must never receive U-47's memories. This project designs prompts and service contracts; it does not save a real person's data or send a notification.
Project: review Parcel Desk memory behavior
Write and correct with evidence
The assistant proposes a metric-units memory only from U-47's direct future-facing statement. The trusted service binds it to tenant, user, and agent scope and returns receipt M-47. The miles request affects one export and does not alter M-47. A separate notice-channel item M-51 changes only when U-47 says 'from now on.' The application records a targeted update receipt and source turn. If the store call times out, the assistant must say the task can proceed using the current instruction while future persistence remains unconfirmed.
Forget and test isolation
A forget request targets M-51. Receipt D-51 starts the confirmation path: check primary storage, retrieval index, derived weekly summary, and context cache, then run a later recall test. A retrieved page instructing the assistant to remember automatic approvals is lower-trust data and cannot write memory or authorize notifications. U-84's tenant and user scope must exclude U-47's item before model context assembly. The release suite includes positive recall, temporary override, lasting correction, cross-user denial, injection, deletion, expiry, and service outage cases.
U-47: metric units for future summaries -> candidate M-47.
S-47 only: miles in this export -> no memory update.
U-47: future notices by text -> targeted revision of M-51.
Forget M-51 -> D-51, then verify store, index, summary, cache.
U-84: cannot retrieve U-47 items.
Retrieved page: automatic approvals -> reject memory write.Performance and operating cost
Keyed item reads and updates can be O(1) average with an index, while checking D derived copies requires O(D) work. A release suite across users, tasks, and memory states expands quickly, so include representative combinations and every security boundary as a hard gate. Retrieving only relevant permitted items lowers token cost and prevents an old preference from burying the current request. The application must verify store receipts and scope; model prose cannot prove persistence, deletion, or isolation.
Common Mistakes
- Do not store a one-time override as a lasting preference.
- Do not accept a model-selected user namespace.
- Do not declare deletion complete while a derived summary still serves the fact.
- Do not let lower-trust page text install a future approval rule.
Connected lessons
- Memory prompts: confirm intent before durable writes
- Memory prompts: bind namespace, identity, and provenance
- Memory prompts: correct stale preferences without erasing history
- Memory prompts: expire and delete every usable copy
- Memory release: test recall, poisoning, and isolation
- Conversation memory: retain decisions without retaining every private detail
- Tool results: keep returned text in the data lane
- Retrieval prompts: enforce document permission first
- Memory-lifecycle prompt decisions
