Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Interruptible compute: price recovery work, not only cheap node hours

Last updated: 5 Oct 20266 min read
tutorial
AdvancedBy AITrove Editorial

Interruptible capacity can be withdrawn by its provider, and any warning window is provider-specific and may be too short for a large checkpoint. A PodDisruptionBudget governs voluntary evictions through the Eviction API; it does not block an involuntary node loss. A discounted node hour has little value if interruption causes duplicate side effects, a long replay, or an SLO breach.

Operational decision

A document-index worker consumes queued archive chunks. Put only replayable workers on an interruptible pool, keep the authoritative job state and checkpoint outside the node, and acknowledge a queue lease only after a durable checkpoint. Bound one work unit so it can finish or safely resume within the available lease and interruption budget. Inject a node loss with no warning in a disposable environment, then verify that another worker resumes at the last committed offset, does not publish a duplicate index generation, and eventually releases the old claim. The checklist below separates provider notice from application correctness. Vary node type and zone eligibility so the pool can obtain replacement capacity; a design that requires one scarce shape may be unavailable precisely when interrupted. Keep a non-interruptible fallback for critical deadlines and define when it is activated, rather than assuming the discounted pool can always replace itself.

Output
Document-index interruption acceptance
Authoritative state: durable checkpoint and idempotent generation key
No-notice test: worker killed during publish and replayed once
Lease rule: acknowledgement follows durable checkpoint
Replacement: compatible capacity available outside one node shape
Deadline: fallback pool starts before queue age violates the service objective
Cost: node saving minus replay, fallback, and transfer cost

Cost and verification

The relevant unit is cost per completed indexed archive, including retried CPU, queue retention, fallback nodes, and data transfer. Measure interruption frequency, lost work per event, replay duration, duplicate-publish attempts, and queue age. A short checkpoint interval raises storage writes; a long interval raises replay cost. Find the interval from observed work-unit cost and recovery objective rather than copying a provider notice duration into the application design.

Common Mistakes

  • Do not treat a PDB as protection from provider interruption.
  • Do not move a non-idempotent publisher onto interruptible nodes without a replay contract.
  • Do not count a discount without charging replay and fallback capacity to the same workload.

Connected lessons

Practice and check

devops
operations
Storage details