Independently trained networks can disagree on a case even when their averaged class probability looks confident; both signals need a review policy.
Deep ensembles, seed diversity and predictive disagreement
Train independent members
A deep ensemble uses separately initialized models trained on the same declared training set, ideally with independently shuffled batches and documented seeds. Five checkpoints from one training run are not five independent fits. Keep architecture, class mapping, preprocessing and evaluation split consistent across members. For receipt damage triage, group receipts by store or acquisition session before any split so shared print defects do not leak into final evaluation. The split lesson applies to every member.
Average probabilities, not labels
Each member produces a probability distribution over the same classes. Average those distributions, then apply the operational threshold to the average. Majority vote discards score strength; averaging raw logits changes scale and is not generally equivalent. The code computes predictive entropy of the mean and the average member entropy. Their difference is a useful disagreement signal, but neither number alone proves calibrated risk.
Separate uncertainty questions
High mean predictive entropy can arise because every model is unsure or because confident members disagree. The disagreement gap helps distinguish these patterns, yet it is sensitive to member diversity and calibration. Inspect both along with error rate. A severe defect that all models confidently miss will have low disagreement, so no uncertainty score replaces a stress test on uncommon damage and new scanners. Selective risk turns a chosen score into a review threshold.
Account for cost and correlation
Training M independent members roughly multiplies training compute by M and stores M sets of weights. At serving, batching can share hardware overhead, but inference work and resident memory still grow. If all members share the same label errors or data blind spots, extra models add cost without reliable uncertainty. Compare ensemble performance with a single well-tuned model and record per-member failure overlap by store, lighting and defect type.
Evaluate review yield
Select the review threshold on a development set. On untouched data, report severe-defect recall, risk among auto-approved cases, human review volume and error capture by the uncertainty queue. The project checks those quantities under scanner and store shift. Do not report only an averaged accuracy score while operators face an unbounded queue.
Implementation
from math import log
def entropy(probabilities):
if any(value < 0 for value in probabilities) or abs(sum(probabilities) - 1) > 1e-9:
raise ValueError("class probabilities must form a distribution")
return -sum(value * log(value) for value in probabilities if value > 0)
member_probabilities = [(.73, .18, .09), (.61, .28, .11), (.24, .65, .11)]
class_count = len(member_probabilities[0])
mean_probability = tuple(sum(member[class_id] for member in member_probabilities)
/ len(member_probabilities) for class_id in range(class_count))
predictive_entropy = entropy(mean_probability)
mean_member_entropy = sum(map(entropy, member_probabilities)) / len(member_probabilities)
disagreement = predictive_entropy - mean_member_entropy
assert abs(sum(mean_probability) - 1) < 1e-9
assert disagreement >= -1e-12
assert predictive_entropy > mean_member_entropyPerformance and operating cost
Aggregating M models across C classes costs O(MC) time and O(C) streaming accumulator space once their outputs exist. Model inference dominates: roughly M forward passes and M parameter sets, with possible parallelism at increased memory use. The entropy calculation is cheap but requires calibrated, aligned class probabilities. The sample distributions illustrate disagreement arithmetic, not a measured error probability.
Common Mistakes
- Do not call several checkpoints from one run independent ensemble members.
- Do not use low disagreement as proof that a prediction is correct.
- Do not choose the review threshold on the final test set.
