A backup is a recovery promise, not merely a file written by a scheduled job. Define how much data loss is tolerable and how long the service may be unavailable, then select backup frequency and restore procedures accordingly. Protect the copy from the same credentials and failure domain as the live database. A restore drill uses an isolated environment to verify that the file is readable, schema migrations are compatible, and a known case can be queried. Record the elapsed time and any manual steps. Do not test a restore by overwriting production data.
Backup and Restore Drills
Working case
A nightly case database snapshot exists, but no one has restored it. During a drill, the team restores it to a separate test database, runs the application migration compatibility checks, and confirms case 47 plus its inspection 93 are present. The measured restore takes 42 minutes, longer than the team's 30-minute recovery target. That gap prompts a change in backup design or recovery procedure. Merely seeing a successful backup job in a dashboard would not have revealed the delay or a missing encryption key.
Implementation
set -euo pipefail
pg_dump --format=custom --file="${BACKUP_OUTPUT_PATH}" "${SOURCE_DATABASE_URL}"
pg_restore --list "${BACKUP_OUTPUT_PATH}" > "${BACKUP_CATALOG_PATH}"
# Restore only into an isolated test database with separate credentials.
pg_restore --dbname="${RESTORE_TEST_DATABASE_URL}" "${BACKUP_OUTPUT_PATH}"Cost and boundaries
A full snapshot transfers O(B) bytes for database size B and consumes comparable protected storage, with additional space for retention. Frequent backups reduce potential data loss but increase transfer and storage cost. Incremental approaches reduce repeated bytes while making restore chains more complex. A drill consumes isolated compute time but turns an assumption into evidence. Test credential access, decryption, schema version, and application reads as part of recovery; a raw database restore alone may not restore the service.
Common Mistakes
- Do not treat backup success as proof of restore success.
- Do not store the only backup under the same failure and access boundary as production.
- Do not run a restore drill against the live database.
Connected lessons
Deployment and Scale; Environment Configuration and Secret Boundaries; Continuous Integration and Release Gates; Horizontal Scale and Shared State; Release checks: prove the critical route and prepare a rollback; Observability: connect user failure to a safe request trace; DevOps: delivery, infrastructure, and reliable operations.
Failure trace
A backup job has run every night, but the only restore test uses an empty database. When a real incident occurs, the restored data lacks recent attachments and the application cannot start against the recovered schema. Rehearse a full restore into an isolated environment with representative records and media, then request key user flows and measure both data age and recovery time.
Verification
- Restore into an isolated target and compare counts, selected IDs, and attachment checksums.
- Start the application against the restored schema and perform one authenticated read and write.
- Record the actual recovery time and newest restored record, then compare with the service promise.
Decision note
A backup artifact is evidence of copying bytes, not proof of recoverability. Keep restore credentials, retention, and runbook steps tested together; protect the drill environment from accidentally sending real notifications.
Apply and check
Build Project: release and recovery drill for a content service; then check the boundary with Web Development: offline and delivery contracts quiz.
Further connections
Backup Restore with Deletion Tombstones.
