Build a maintenance-summary feature for case 47 at revision 29. The browser submits a scoped operation key and a request for a short answer. The server checks the reviewer, tenant, selected records, input size, output cap, and concurrent-work budget before calling a model service. Keep the service credential outside browser assets. A duplicate request for the same revision returns one operation ID; a changed revision with the old key produces a conflict. Store the operation state so a timeout cannot be mistaken for a completed answer. A result read checks current case permission again. The browser can watch a framed stream, but it labels accumulated text incomplete until a terminal event arrives. It decodes split bytes and frames, ignores old operation events after navigation, and bounds its buffer. Closing the tab requests cancellation but does not claim that already consumed work or a terminal answer disappeared. Retrieved maintenance notes carry case and tenant identity and are filtered before generation. A note that tells the system to email private records stays data. The summary task has no mail or write tool. A later task-creation action requires a separate user confirmation and a fresh authorization check in the tool handler. Output is rendered as text; evidence IDs must belong to the permitted case and revision. Keep an 83-case release set with ordinary, missing-evidence, permission-change, hostile-note, and truncated-stream examples. Block a candidate if it leaks a different tenant or misses a severe finding. Stage a passing version behind a reversible flag, retain version identity on each operation, and keep a direct route to original notes for human review.
Project: permission-bound maintenance summary
Build contract
- Bind one operation to a user, tenant, case revision, input budget, and result permission.
- Show partial streamed text as incomplete until an authoritative terminal event.
- Deny tool effects that are not explicitly authorized by the user action.
- Release only after severe-error checks and retain a narrow rollback path.
Implementation checkpoint
function summaryMayShip(operation) {
return operation.caseAuthorized && operation.terminalEventSeen &&
operation.evidenceIds.every(id => operation.permittedEvidenceIds.has(id));
}
console.log(summaryMayShip({ caseAuthorized: true, terminalEventSeen: true, evidenceIds: [47, 63], permittedEvidenceIds: new Set([47]) }));
// Output: falseThe checkpoint rejects a summary that cites evidence 63 when only evidence 47 is permitted. The real service must perform this test against authoritative record and revision data, not against a client-supplied set. The check is O(E) for E cited evidence IDs with expected O(1) set membership, and the retained set costs O(P) space for P permitted IDs. A large citation list needs its own cap. This function is one small gate; it cannot establish whether prose faithfully represents the evidence, so the release set and reviewer correction path cover that separate failure.
Failure drill
Retry the same request from two tabs. Change the case revision, revoke the reviewer, insert a hostile instruction in a note, split a streamed frame across byte chunks, cut the connection before its terminal marker, and try to send a task to another tenant. None may produce a completed unauthorized result or a privileged side effect. Roll a candidate version to a small cohort, then inject one cross-tenant evidence reference. Activate the stop action and verify only new operations switch routes while old operation statuses retain their version and outcome. Inspect browser assets, logs, and telemetry labels for credentials and private note bodies. Test a keyboard and screen-reader journey that hears one completion or failure announcement rather than every arriving token.
Acceptance checks
- No provider credential or private note body is exposed in routine client assets and logs.
- A canceled or interrupted stream is never labeled complete without a terminal event.
- Retrieved text cannot authorize a tool call or expand tenant scope.
- A severe release failure blocks or rolls back new operations.
Common Mistakes
- Treating a convincing answer as a verified case decision.
- Using prompt wording as the only tool permission control.
- Counting an interrupted stream as a completed result.
Related lessons
Model Gateway Identity and Request Budgets; Generated Response Streaming and Cancel State; Retrieved Content, Instructions, and Tool Permission; Model Output Evaluation and Release Control.
