Skip to content
AITroveRead. Build. Understand.
Make this comfortable

CronJob schedule safety: missed runs, overlap, and replay

Last updated: 5 Oct 20266 min read
tutorial
AdvancedBy AITrove Editorial

A Kubernetes CronJob creates Jobs according to a schedule. concurrencyPolicy governs overlap among Jobs created by that same CronJob: Allow runs concurrently, Forbid skips an overlapping start, and Replace terminates the earlier Job. startingDeadlineSeconds limits how late a missed schedule can start. None of those fields turns an external write into an exactly-once operation.

Operational decision

A nightly royalty ledger close runs at 02:17 UTC. Forbid overlap, allow a bounded late start, and give each business date a unique ledger-close key checked by the database. Record the date in durable storage before final publication so a controller restart or manual replay cannot produce a second payout. The fragment assumes the image knows which date to close and has least-privilege credentials supplied outside the manifest. Alert on missed schedules and failed Jobs; a skipped run under Forbid is a business event, not proof that work completed. Test a run that lasts past the next schedule, a controller outage, and a manual retry. Use a time zone explicitly where supported and document daylight-saving behavior if local time matters.

yaml
apiVersion: batch/v1
kind: CronJob
metadata: {name: royalty-close, namespace: finance}
spec:
  schedule: '17 2 * * *'
  timeZone: Etc/UTC
  concurrencyPolicy: Forbid
  startingDeadlineSeconds: 1200
  jobTemplate:
    spec:
      backoffLimit: 2
      template:
        spec:
          restartPolicy: Never
          containers:
            - name: royalty-close
              image: registry.internal/royalty-close@sha256:REPLACE_WITH_FULL_DIGEST

Cost and verification

Each Job creates Pods and consumes requested resources even when the business operation later finds nothing to do. Retained Job histories also consume API and log storage; set retention after deciding how failures will be investigated. A narrow deadline can discard important work after a controller outage, while an unbounded catch-up can flood downstream systems. Test the chosen schedule against real processing duration and recovery requirements.

Common Mistakes

  • Do not rely on Forbid as a database lock across separate CronJobs.
  • Do not assume every scheduled time produces exactly one Job.
  • Do not let a replay pay the same ledger date twice.

Connected lessons

Serverless operating-boundary follow-up

systemd operating follow-up

devops
resilience
Storage details