Temporal validation freezes a model on earlier cases and tests it on later case arrivals whose required outcome horizon has matured.
Temporal validation of a support-resolution model
Split by case creation, not by rows
A start-stop table can hold several rows per case. Randomly splitting those rows puts one case in both training and test data. Split by the immutable case creation date and keep all history for one case together. A realistic release might train on January through March arrivals, tune on April, and hold May for final review. The exact months are a design choice; record them before tuning. Interval rows remain grouped by case ID.
Let outcomes mature
For a nine-day resolution target, a case created one day before the data cutoff cannot yet have a known nine-day open status. Exclude it from the simple fully observed score or handle its censoring through a declared method. The code checks that a test case had at least nine days available between creation and extraction. It does not pretend that this administrative maturity check resolves other causes of early loss. Censor-aware scoring covers observed dropout.
Freeze the entire pipeline
Feature definitions, imputation values, coefficient selection, probability recalibration and censoring-model settings belong to training or tuning data. Later test labels may be used once to report performance, not to choose a new horizon, threshold or subgroup definition. Store a model version and data snapshot hash with predictions. If the live service lacks a historically reconstructed feature, compare only the serving-compatible version.
Read performance with context
Compare case mix, queue policy, event coding, inspection cadence and missing-feature rates across periods. A score drop may be a true model failure, a change in censoring, or a changed outcome definition. Report calibration, Brier score and ranking on the same eligible population and horizon, with counts. A model can look stable overall while overpredicting resolution for a small queue; show group support rather than many unsupported slices.
Write a release decision
A later cohort supplies stronger evidence of transfer than a random split from one month, but it still cannot guarantee future stability. State when reevaluation is due and which workflow changes trigger it. Withhold an automated action if the target, prediction-time features or observation support changed materially. The review project assembles the checks into one packet.
Implementation
from datetime import date
def temporal_holdout(case_records, training_end, test_start, extraction_day,
horizon_days):
if not training_end < test_start <= extraction_day or horizon_days <= 0:
raise ValueError("invalid calendar design")
if len({case_id for case_id, _ in case_records}) != len(case_records):
raise ValueError("duplicate case ID")
training, test, immature = [], [], []
for case_id, created_day in case_records:
if created_day > extraction_day:
raise ValueError("case created after extraction")
if created_day <= training_end:
training.append(case_id)
elif created_day >= test_start:
if (extraction_day - created_day).days >= horizon_days:
test.append(case_id)
else:
immature.append(case_id)
return training, test, immature
cases = [("R601", date(2026, 3, 22)), ("R602", date(2026, 5, 4)),
("R603", date(2026, 5, 28))]
train_ids, test_ids, immature_ids = temporal_holdout(
cases, date(2026, 3, 31), date(2026, 5, 1), date(2026, 5, 31), 9)
assert train_ids == ["R601"]
assert test_ids == ["R602"]
assert immature_ids == ["R603"]Performance and operating cost
Partitioning N case records costs O(N) expected time and O(N) output space. Model fitting and scoring add their own costs. Maturity filtering changes the evaluated arrival population near the extraction edge, so counts and dates must accompany any performance figure.
Common Mistakes
- Do not split interval rows from the same case across periods.
- Do not use the final test month to pick the best feature set.
- Do not call a case mature merely because its outcome happened early while ignoring comparable open cases.
