Release tests, adversarial inputs, cost, privacy, and language parity. Start with the decision you need to make, then use the linked material to build a checkable result.
Learning path
- Prompt injection: test untrusted content at every boundary
- Evaluation sets: measure the failure cases that matter
- Model judges: calibrate rubrics and swap candidate order
- Prompt releases: version the whole decision path and keep a rollback
- Prompt budgets: trade output quality against cost and tail latency
- Prompt privacy: send only the fields needed for the task
- Multilingual prompts: test policy meaning across languages
Connect the path
The foundations define the task and trust boundary. Later modules test whether a result should be used, revised, or sent to a person.
Experiments, releases, and review controls
These lessons turn prompt wording into a tested workflow. Apply each rule to a real input boundary, then inspect whether the decision and any side effect still match the task contract.
- Prompt caching: reuse stable context with a versioned expiry
- Metamorphic tests: verify behavior when harmless details change
- Human handoff: preserve evidence and the reason for uncertainty
- Prompt traces: connect an answer to its inputs, checks, and effects
- Model migration: replay contracts before changing providers or versions
- Fairness checks: test equivalent cases across groups and wording
Decision gates and downstream validation
A reliable answer needs the right information, a valid calculation or decision, and a destination that treats generated content as data.
- Selective answers: measure when to abstain
- Evaluation leakage: keep the release test independent
- Generated output: validate again at the destination boundary
Executable evaluation labs
Use these short programs to check evidence references, abstention, judge consistency, and critical release cases. A passing structural check still needs semantic review where the claim depends on a passage.
- Code lab: reject unknown evidence IDs
- Code lab: score answered cases and abstentions
- Code lab: find pairwise judge order changes
- Code lab: block a critical prompt regression
- Prompt evaluation code labs
Tool result and effect controls
Contain instructions in tool data and reconcile ambiguous effects.
- Tool results: keep returned text in the data lane
- Tool effects: reconcile receipts before retrying
- Agent workflow decisions
- Project: recover a tool workflow without duplicate effects
Negative controls
Test cases that should not produce an answer or external effect.
- Negative controls: test the answer that should not be produced
- Project: verify a contract-renewal prompt at the boundary
- Reasoning and evidence checks
Delivery and telemetry
Connect the tested prompt to a merge gate, a complete release bundle, and a trace that preserves decisions without copying private payloads.
- Prompt changes in CI: test the merge candidate
- Prompt release artifacts: version the whole decision path
- Prompt telemetry: measure failures without copying private payloads
- Continuous integration: test the merge candidate
- Immutable artifacts and release provenance
- Observability: join metrics, logs, and traces
Evaluation governance
Test paired outcomes, labels, attack boundaries, live exposure, and case revisions.
- Paired prompt evaluation: count changes, then inspect uncertainty
- Evaluation labels: adjudicate disagreement before scoring a release
- Adversarial case mutations: test the boundary, not a magic phrase
- Online prompt experiments: define exposure and stop rules first
- Evaluation case ledgers: revise labels without erasing history
- Project: govern a claims-assistant evaluation board
- Prompt evaluation governance decisions
Safety decision boundaries
Route requests, refuse narrowly, validate output, and reserve high-impact decisions for review.
- Request risk routing: classify the action before choosing a response
- Refusal contracts: decline the unsafe effect and preserve useful help
- Sensitive output gates: check the rendered answer before release
- System prompts: treat instructions as guidance, not a secret vault
- High-impact advice: route uncertain decisions to qualified review
- Project: build a support safety boundary
- Prompt safety decisions
Release review packet
Bring case-level checks, costs, promotion, and rollback into one decision.
Cross-locale evaluation
Pair semantic cases across languages and compare decisions.
- Cross-locale evaluation: compare decisions, not word-for-word text
- Project: release a multilingual order notice safely
- Multilingual prompt workflow decisions
Accessible output evaluation
Check artifact failures and rendered user tasks before promotion.
- Accessible output evaluations: measure failures by artifact and user task
- Project: release an accessible transit disruption alert
- Accessible prompt output decisions
Browser-agent recovery
Resolve an uncertain submit with a receipt or authoritative state.
- Browser effect recovery: resolve ambiguous submissions before retrying
- Project: reschedule an appointment through a browser safely
- Browser-agent prompt decisions
Voice-agent evaluation
Test event timing, noise, receipts, and final account state.
- Real-time voice evaluations: test timing and state, not transcript fluency
- Project: a voice call that renews the right library item
- Real-time voice prompt decisions
Synthetic evaluation-case engineering
Generate cases from a contract, validate labels, protect privacy, and keep holdouts independent.
- Synthetic evaluation cases: mutate a contract, not a customer's record
- Synthetic case oracles: reject invalid labels before scoring a model
- Synthetic case coverage: count distinct decisions, not rewritten sentences
- Synthetic cases are not automatically private
- Synthetic evaluation holdouts: stop the generator from teaching the test
- Project: build a checked synthetic return-triage suite
- Synthetic evaluation-case decisions
Generated media release review
Check continuity, captions, transcript, and the actual published asset.
- Generated video release: inspect continuity, claims, and alternatives
- Project: release an audio and video guide for a fictional exhibit
- Generated speech and video prompt decisions
Spreadsheet calculation release gate
Recalculate the final workbook and compare it with the reviewed artifact.
- Spreadsheet release checks: recalculate, reconcile, then inspect
- Project: review a depot replenishment workbook
- Spreadsheet prompt and workbook release decisions
Specialist run release gate
Bound timeouts and retries, review the trace, and name missing results.
- Specialist workflows: bound retries and review the full trace
- Project: coordinate a parcel-platform incident review
- Specialist-agent coordination decisions
Research brief release gate
Tie the final claim to the frozen packet, exclusions, calculations, and reviewer.
- Research synthesis release: expose the method and its limits
- Project: synthesize conflicting parcel-scan trials
- Research-synthesis prompt decisions
Tutor state and release checks
Minimize learner data and test new-case transfer before release.
- Tutor release checks: minimize learner state and test transfer
- Project: review a retry-safety training tutor
- Tutoring and assessment prompt decisions
UI release evidence
Compare rendered states and test keyboard, semantics, and truthful feedback.
- UI release prompts: compare appearance and verify access
- Project: implement and review a dispatch queue UI
- UI-to-code prompt decisions
PDF export release
Verify the delivered file's text, pages, structure, and approved claims.
- PDF export prompts: verify the delivered file again
- Project: deliver a North Quay maintenance briefing
- Document and presentation prompt decisions
Meeting recap release
Check access, sensitive material, correction routes, and side-effect receipts.
- Meeting recap release: review scope, access, and corrections
- Project: review a Beacon pilot meeting record
- Meeting-record prompt decisions
Architecture release evidence
Load-test the chosen contract and stage its rollout with rollback ownership.
- Architecture release: require evidence, rollback, and ownership
- Project: review a receipt-ingest architecture decision
- Architecture-decision prompt checks
API docs release evidence
Check links, examples, shared status lists, and deployed behavior.
- API docs release: test links, examples, and live version
- Project: review Shipment Status API documentation
- API-documentation prompt decisions
Infrastructure release evidence
Verify live resources, behavior, drift, and rollback limits after apply.
- Infrastructure release: verify live state and rollback limits
- Project: review an Aster queue retention change
- Infrastructure-change prompt decisions
Incident handoff and closure
Keep an action ledger, verify recovery, and correct earlier claims.
- Incident handoffs: preserve actions and decisions
- Incident resolution: verify recovery and correct the record
- Project: draft Harbor Checkout incident updates
- Incident-communication prompt decisions
Retrieval release gates
Test revocation, deletion, timeouts, quality, and denied-content leakage.
- Retrieval release: test permission changes and deletion
- Project: secure the Meridian maintenance assistant
- Authorized-retrieval prompt decisions
Memory deletion and release
Confirm every usable copy is gone and test persistent attack boundaries.
- Memory prompts: expire and delete every usable copy
- Memory release: test recall, poisoning, and isolation
- Project: review Parcel Desk memory behavior
- Memory-lifecycle prompt decisions
Recurring task controls
Deduplicate retries, bound alerts and effects, and retain run evidence.
- Recurring prompts: survive retries without duplicate effects
- Recurring prompts: notify only on actionable changes
- Recurring prompts: renew authority for sensitive actions
- Recurring prompts: make every run replayable and auditable
- Project: build a North Pier recurring review
- Recurring prompt workflow decisions
Terminal-agent release controls
Interpret exit status, preserve changes, respect permissions, and verify final claims.
- Terminal agents: distinguish exit status from useful output
- Terminal agents: request only the missing execution scope
- Terminal agents: keep the patch within owned files
- Terminal agents: make the final claim match verified checks
- Project: verify a Manifest Gate importer fix
- Terminal-agent prompt decisions
Chart release checks
Store reproducible specs and review rendered and nonvisual outputs.
- Chart prompts: produce a reproducible spec and data lineage
- Chart prompts: review rendered and nonvisual outputs
- Project: review a cold-chain temperature chart
- Chart-creation prompt decisions
Email and calendar effect gates
Review recipients, receipts, event updates, and recurrence scope.
- Email prompts: review To, Cc, Bcc, and reply scope
- Email prompts: separate drafted commitments from sent messages
- Calendar prompts: distinguish create, update, and cancellation
- Calendar prompts: scope recurrence changes to the right instances
- Project: review an Aster vendor reply and invite
- Email and calendar prompt decisions
Route feasibility and release
Compute windows, inspect map points, and gate dispatch effects.
- Location prompts: compute stop-window feasibility before recommending
- Location prompts: review the map and gate dispatch effects
- Project: review a Quartz field-service route
- Location-aware prompt decisions
Qualitative analysis and decisions
Adjudicate labels, retain counterexamples, reconcile people, and gate claims.
- Interview prompts: keep coder disagreement visible
- Interview prompts: build a theme evidence ledger
- Interview prompts: count people, not repeated excerpts
- Interview prompts: turn bounded findings into testable decisions
- Project: analyze Parcel Window pickup interviews
- Qualitative interview prompt decisions
Forecast evaluation and decisions
Test baselines, distinguish scenarios, check intervals, and gate stock actions.
- Forecast prompts: compare against a rolling baseline
- Forecast prompts: keep scenarios separate from predictions
- Forecast prompts: label intervals and check coverage
- Forecast prompts: gate inventory recommendations on evidence
- Project: review Aster depot kit forecasts
- Forecast prompt decisions
Security triage decisions
Apply severity rules, quarantine log instructions, and gate containment.
- Security alert prompts: apply scoped severity and escalation rules
- Security alert prompts: quarantine instructions embedded in logs
- Security alert prompts: gate containment and verify receipts
- Project: triage Orion authentication alerts
- Security-alert triage decisions
Survey testing and release
Pilot the form, reconcile counts, protect open text, and bound claims.
- Survey prompts: pilot the instrument before fielding
- Survey prompts: reconcile invitations, submissions, and item denominators
- Survey prompts: code open text without exposing respondents
- Survey prompts: release findings with explicit scope
- Project: review Meridian pickup survey findings
- Survey prompt decisions
Support decisions and handoff
Verify permissions, ask focused questions, and separate draft from effect.
- Support prompts: verify entitlement and requester authority
- Support prompts: ask only the question that changes the next step
- Support prompts: separate reply, handoff, and resolution
- Project: handle Aster seat activation ticket T-47
- Customer-support case decisions
Accessibility review review
Verify state, arithmetic, and release boundaries.
- Accessibility prompts: connect form errors and status changes
- Accessibility prompts: measure contrast and describe image purpose
- Accessibility prompts: inspect zoom and reflow states
- Accessibility prompts: write reproducible findings and retest fixes
- Project: review Harbor booking access across states
- Accessibility review prompt decisions
Product analytics review
Verify state, arithmetic, and release boundaries.
- Analytics prompts: pin cohort windows and time zones
- Analytics prompts: review read-only queries and totals
- Analytics prompts: explain segment gaps without inventing causes
- Analytics prompts: turn a metric packet into a bounded decision
- Project: reconcile Meridian's signup funnel
- Product analytics prompt decisions
Localization QA review
Verify state, arithmetic, and release boundaries.
- Localization prompts: render typed dates, numbers, and units
- Localization prompts: reconcile terminology and tone
- Localization prompts: use pseudolocale and real layout checks
- Localization prompts: gate release on catalog and rendered retests
- Project: review Harbor's localized delivery notices
- Localization QA prompt decisions
Product experiments review
Verify state, arithmetic, and release boundaries.
- Experiment prompts: reconcile missing outcomes by assigned unit
- Experiment prompts: separate effect size from uncertainty
- Experiment prompts: investigate slices and guardrail movement
- Experiment prompts: gate rollout on a reproducible packet
- Project: review Aster's signup experiment
- Product experiment prompt decisions
Invoice exception review review
Verify comparisons and release boundaries.
- Invoice prompts: match order, receipt, and bill
- Invoice prompts: distinguish duplicate ingests from credits
- Invoice prompts: route exceptions before approval
- Project: reconcile a North Quay invoice exception
- Invoice exception prompt decisions
Database backfill review
Verify comparisons and release boundaries.
- Backfill prompts: account for writes during migration
- Backfill prompts: compare count and field-level parity
- Backfill prompts: hold cutover until fallback is rehearsed
- Project: review Harbor's order-status backfill
- Database backfill prompt decisions
Search relevance evaluation review
Verify comparisons and release boundaries.
- Search prompts: calculate ranking metrics on judged pairs
- Search prompts: interpret clicks and gate a ranking release
- Project: review Beacon Supply search relevance
- Search relevance prompt decisions
Annotation quality workflow review
Verify comparisons and release boundaries.
