A scheduled function can outlive its interval, fail after partial work, or be invoked again while a prior run remains active. A process-local mutex cannot coordinate separate execution environments. Use a durable lease or compare-and-set marker with an expiry, and make each unit of work restartable. A run record should distinguish scheduled time, actual start, last checkpoint, completion, and repair owner. If missed intervals must be replayed, define the catch-up window and a backfill rate that does not overwhelm the source system.
Scheduled functions: control overlap, catch-up, and missed runs
Operational decision
A nightly entitlement reconciliation usually takes 11 minutes but occasionally takes 39. The next trigger must not start another full scan over the same accounts. Acquire a lease for the intended interval, checkpoint each account range, and renew only while progress continues. Kill the worker after range 17, then run the next trigger and show that it resumes the unfinished interval without applying duplicate entitlement changes. Expire a stalled lease only after its fencing token prevents the old holder from committing. Record skipped or delayed intervals and alert when the oldest unfinished interval exceeds the reporting deadline. A manual rerun uses the same interval identity and audit trail.
Entitlement reconciliation
Interval ID: UTC scheduled date
Lease: owner, expiry, fencing generation
Progress: last committed account range
Retry: resume incomplete interval
Catch-up: bounded oldest unfinished interval
Completion: count and reconciliation checksumCost and verification
Scanning A accounts is O(A) per completed interval. Replaying a full scan after every failure multiplies provider calls and can create correlated rate-limit failures. Checkpointing adds writes and operational complexity but caps repeated work. Measure run delay, duplicate attempts, lease loss, backlog age, and the time until the entitlement view is consistent.
Common Mistakes
- Do not use a process-local lock for a distributed schedule.
- Do not let an expired worker commit after a new lease holder starts.
- Do not mark the scheduled trigger as proof of completed work.
Connected lessons
- DevOps: delivery, infrastructure, and reliable operations
- CronJob schedule safety: missed runs, overlap, and replay
- Leader leases: fence side effects after ownership changes
- Database backfills: checkpoint progress without racing live writes
Practice and check
- Project: prove a serverless release survives duplicate events and rollback
- DevOps serverless operations decisions
