Model migration should preserve the task contract, accepted output schema, tool permissions, and evaluation cases while testing the new model's behavior. Capability names and parameter ranges can differ, so translate configuration deliberately instead of copying an old request object. Run both models on the same redacted cases and compare business outcomes, abstention, latency, cost, and critical slices. Record the model identifier and migration date with each output; a shared prompt string does not prove equivalent decisions.
Model migration: replay contracts before changing providers or versions
Decision in practice
A claims team moves its triage workflow to a newer model. The old prompt uses a partial assistant prefill that the new endpoint does not accept. The team removes that mechanism, uses the supported structured-output option, and runs 96 held-out claims through both routes. The new route produces fewer format errors but more unsupported approvals in the missing-receipt slice. Migration stops until the decision boundary is fixed and a fresh holdout confirms the repair. The old route remains available for rollback during the canary.
Frozen contract: labels, evidence IDs, unknown state, no direct write.
Compare: 96 held-out cases on old and candidate models.
Gate: zero unsupported approvals in missing-receipt slice.
Record: model ID, prompt hash, schema version, settings, result IDs.Performance and operating cost
Dual running roughly doubles model-call cost during evaluation and can increase temporary storage. That expense is bounded by the migration window and smaller than an unobserved production regression. Keep a stable test set plus fresh cases to detect prompt overfitting. If the provider changes token accounting or tool semantics, measure actual billed usage and effect logs rather than extrapolating from old settings. Release the candidate gradually and keep a compatible rollback bundle.
Common Mistakes
- Do not assume the same prompt gives equivalent behavior across models.
- Do not copy unsupported parameter names blindly.
- Do not declare migration complete from a format score alone.
Connected lessons
- Prompt Engineering
- Production prompt engineering
- Model settings: change one generation variable against a fixed case set
- Prompt releases: version the whole decision path and keep a rollback
- Prompt traces: connect an answer to its inputs, checks, and effects
- Fairness checks: test equivalent cases across groups and wording
- Project: regression-test a customer triage prompt
- Advanced prompt engineering decisions
Related implementation
Continue with: API changes: explain client impact and migration.
