Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Spring Batch partition restart: preserve worker names and input slices

Last updated: 5 Oct 20264 min read
tutorial
IntermediateBy AITrove Editorial

A failed partition should resume its recorded slice; changing the grid hint or source boundaries during restart can create a different workload.

The repository records worker executions

A partition manager creates named worker StepExecutions with their input contexts. On restart, a splitter can reconstitute failed work from those records rather than recompute an arbitrary new distribution. Treat the original partition names, bounds and manifest version as part of the recovery contract. Range ownership defines what each worker may write.

A grid is capacity, not identity

The requested grid size is a hint for the initial division, and a restart may use the saved partition set. Do not assume changing it from four to seven repartitions an interrupted job safely. If a deployment requires a new split, launch a new job instance against an explicitly versioned source after reconciling committed output. Keep destination writes keyed by manifest and receipt reference, so a repeated partition cannot duplicate side effects. Writer idempotency is the second layer of protection.

Run the hard-stop case

Kill the worker process after one partition commits while another has not. Inspect the repository for each partition name, status and context. Restart with the same job identity and verify only incomplete slices run; assert the final destination set equals the source once. A unit test of partitioner arithmetic is necessary but cannot prove repository-driven restart.

Implementation contract

Java
for (Map.Entry<String, ExecutionContext> partition : persistedRanges.entrySet()) {
    String partitionName = partition.getKey();
    int start = partition.getValue().getInt("receipt.startInclusive");
    int end = partition.getValue().getInt("receipt.endExclusive");
    assertThat(partitionName).startsWith("receiptRange-");
    assertThat(start).isLessThan(end);
}

Cost and verification

Persisted partition metadata costs rows and coordination. Recovery can save substantial replay, but operational checks must track each worker rather than just the manager status.

Common Mistakes

  • Do not treat a new grid size as permission to redistribute an incomplete job.
  • Do not delete failed worker metadata before reconciliation.
  • Do not count manager completion as proof of every expected receipt key.

Read next

Spring Batch partitioner: assign disjoint receipt ranges, Spring Batch restart tests: assert metadata rows, not repeated ID values, Spring Batch killed worker: fence it, recover STARTED, then restart, Spring Batch writers: use a stable source key when a chunk is replayed, Spring Batch ExecutionContext keys: give each stream its own checkpoint.

spring
spring-batch
batch-partition-restart-shape
Storage details