A Kubernetes CronJob creates Jobs according to a schedule. concurrencyPolicy governs overlap among Jobs created by that same CronJob: Allow runs concurrently, Forbid skips an overlapping start, and Replace terminates the earlier Job. startingDeadlineSeconds limits how late a missed schedule can start. None of those fields turns an external write into an exactly-once operation.
CronJob schedule safety: missed runs, overlap, and replay
Operational decision
A nightly royalty ledger close runs at 02:17 UTC. Forbid overlap, allow a bounded late start, and give each business date a unique ledger-close key checked by the database. Record the date in durable storage before final publication so a controller restart or manual replay cannot produce a second payout. The fragment assumes the image knows which date to close and has least-privilege credentials supplied outside the manifest. Alert on missed schedules and failed Jobs; a skipped run under Forbid is a business event, not proof that work completed. Test a run that lasts past the next schedule, a controller outage, and a manual retry. Use a time zone explicitly where supported and document daylight-saving behavior if local time matters.
apiVersion: batch/v1
kind: CronJob
metadata: {name: royalty-close, namespace: finance}
spec:
schedule: '17 2 * * *'
timeZone: Etc/UTC
concurrencyPolicy: Forbid
startingDeadlineSeconds: 1200
jobTemplate:
spec:
backoffLimit: 2
template:
spec:
restartPolicy: Never
containers:
- name: royalty-close
image: registry.internal/royalty-close@sha256:REPLACE_WITH_FULL_DIGESTCost and verification
Each Job creates Pods and consumes requested resources even when the business operation later finds nothing to do. Retained Job histories also consume API and log storage; set retention after deciding how failures will be investigated. A narrow deadline can discard important work after a controller outage, while an unbounded catch-up can flood downstream systems. Test the chosen schedule against real processing duration and recovery requirements.
Common Mistakes
- Do not rely on Forbid as a database lock across separate CronJobs.
- Do not assume every scheduled time produces exactly one Job.
- Do not let a replay pay the same ledger date twice.
Connected lessons
- DevOps: delivery, infrastructure, and reliable operations
- Queue consumers: acknowledgement, idempotency, and backlog
- Database change safety: expand, migrate, contract
- Alert design: page on impact and include a first action
