A failed partition should resume its recorded slice; changing the grid hint or source boundaries during restart can create a different workload.
Spring Batch partition restart: preserve worker names and input slices
The repository records worker executions
A partition manager creates named worker StepExecutions with their input contexts. On restart, a splitter can reconstitute failed work from those records rather than recompute an arbitrary new distribution. Treat the original partition names, bounds and manifest version as part of the recovery contract. Range ownership defines what each worker may write.
A grid is capacity, not identity
The requested grid size is a hint for the initial division, and a restart may use the saved partition set. Do not assume changing it from four to seven repartitions an interrupted job safely. If a deployment requires a new split, launch a new job instance against an explicitly versioned source after reconciling committed output. Keep destination writes keyed by manifest and receipt reference, so a repeated partition cannot duplicate side effects. Writer idempotency is the second layer of protection.
Run the hard-stop case
Kill the worker process after one partition commits while another has not. Inspect the repository for each partition name, status and context. Restart with the same job identity and verify only incomplete slices run; assert the final destination set equals the source once. A unit test of partitioner arithmetic is necessary but cannot prove repository-driven restart.
Implementation contract
for (Map.Entry<String, ExecutionContext> partition : persistedRanges.entrySet()) {
String partitionName = partition.getKey();
int start = partition.getValue().getInt("receipt.startInclusive");
int end = partition.getValue().getInt("receipt.endExclusive");
assertThat(partitionName).startsWith("receiptRange-");
assertThat(start).isLessThan(end);
}Cost and verification
Persisted partition metadata costs rows and coordination. Recovery can save substantial replay, but operational checks must track each worker rather than just the manager status.
Common Mistakes
- Do not treat a new grid size as permission to redistribute an incomplete job.
- Do not delete failed worker metadata before reconciliation.
- Do not count manager completion as proof of every expected receipt key.
Read next
Spring Batch partitioner: assign disjoint receipt ranges, Spring Batch restart tests: assert metadata rows, not repeated ID values, Spring Batch killed worker: fence it, recover STARTED, then restart, Spring Batch writers: use a stable source key when a chunk is replayed, Spring Batch ExecutionContext keys: give each stream its own checkpoint.
