A run audit connects the scheduler event to the data snapshot, prompt and model version, validation results, output artifact, notification decision, and any effect receipts. Store identifiers and hashes where full payload retention is unnecessary. Replaying a run means using the same approved input snapshot and versions in a controlled environment; it does not guarantee identical model wording. Compare decision fields and critical claims against a reference oracle. Keep failed and skipped runs, since excluding them creates a misleading success rate. Separate an attempted run from a verified completed outcome.
Recurring prompts: make every run replayable and auditable
Operational case
The manager disputes an alert about parcel H-71. The team retrieves run identity W-47, snapshot hash S-47, prompt version review-v3, validation status, and alert receipt A-71. A replay on that frozen snapshot reproduces the decision to flag a new hold, while a later live query shows the parcel was released after the original cutoff. The team corrects the present-day state without rewriting the historical report. Another run that failed to fetch partition P-4 remains visible as incomplete, not counted as a successful no-change run.
Run W-47: schedule event, window, snapshot S-47, prompt review-v3.
Validation: all expected partitions present.
Decision: new hold H-71; alert receipt A-71.
Replay: frozen S-47, compare decision fields.
Later release: new fact, not a rewrite of W-47.
Failed P-4 run: incomplete and visible.Performance and operating cost
Writing a compact ledger row per run is O(1); retaining full snapshots can be O(NR) over N records and R runs, so use retention limits and content-addressed storage where appropriate. Replay costs another model call and possibly validation work. It is most useful for disputed decisions and release regressions, not routine unchanged checks. Count failed, skipped, and duplicate attempts explicitly when reporting reliability.
Common Mistakes
- Do not claim byte-for-byte model reproduction from a replay.
- Do not remove failed runs from the denominator of reliability metrics.
- Do not overwrite old evidence when a later status changes.
Connected lessons
- Production prompt engineering
- Prompt Engineering
- Prompt traces: connect an answer to its inputs, checks, and effects
- Evaluation case ledgers: revise labels without erasing history
- Prompt release artifacts: version the whole decision path
- Recurring prompts: define the trigger and stop condition
- Recurring prompts: bind local time and data windows
- Recurring prompts: validate fresh inputs and empty states
- Recurring prompts: survive retries without duplicate effects
- Recurring prompts: notify only on actionable changes
- Recurring prompts: renew authority for sensitive actions
- Project: build a North Pier recurring review
- Recurring prompt workflow decisions
