Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Design a validation subsample for outcome labels

Last updated: 5 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

An outcome-label validation subsample pairs the production label with a more reliable adjudicated label on selected records so classification error can be estimated.

Define what the two labels mean

A support system flags a ticket as a deadline breach from event logs. An audit team checks the ticket history and policy clock to assign an adjudicated breach label. The latter must follow a written rubric and should be reviewed without knowledge of the desired policy effect. It is a reference process, not a magical error-free truth. Record disagreement reasons, ambiguous cases and adjudicator agreement. The proxy-label lesson explains how operational labels drift from the intended outcome.

Sample across the cells that matter

A random sample of tickets may contain too few production-positive cases to estimate false positives or enough tickets from only the largest branch. Stratify by production label and by the policy group or period whose rates will be compared. Fix an inclusion probability for each cell and retain it with every sampled ticket. Oversampling rare flagged tickets is useful for estimating error rates, but raw validation-sample prevalence then does not represent the full population.

Keep review independent of downstream effects

Do not choose cases because their adjudicated label would make the policy result favorable. Sample IDs before reviewers inspect the ticket. Blind reviewers to treatment or month when practical, while retaining enough information to adjudicate the deadline accurately. A reference process that uses the production flag as its main evidence can reproduce the same mistake. The error-rate lesson needs paired labels from this step.

Track selection and missing review

A sampled case may be impossible to adjudicate because the audit log is gone. Count such failures by stratum. If unresolved records are concentrated among flagged tickets after the policy launch, complete-case error rates may be biased. The code selects a reproducible number of IDs per production-label and branch cell; it does not solve nonresponse or calculate uncertainty. Missingness sensitivity provides a separate tool.

Preserve the sampling ledger

Store the frame count, selected count, selection rule, random seed and review status for every cell. To estimate a population quantity from an unequal-probability audit, use the design weights or a suitable likelihood. Do not discard the cell counts after adjudication. Design weighting covers the general principle; the later correction lesson uses conditional error rates rather than a naive audit prevalence.

Implementation

python
from random import Random

def draw_validation_ids(ticket_records, per_cell, seed):
    cells = {}
    for ticket_id, branch_group, production_label in ticket_records:
        if production_label not in (0, 1):
            raise ValueError("production label must be binary")
        cells.setdefault((branch_group, production_label), []).append(ticket_id)
    randomizer = Random(seed)
    selection = []
    for cell, ticket_ids in sorted(cells.items()):
        if len(ticket_ids) < per_cell:
            raise ValueError("cell too small: " + str(cell))
        chosen = randomizer.sample(sorted(ticket_ids), per_cell)
        inclusion_probability = per_cell / len(ticket_ids)
        selection.extend((ticket_id, cell, inclusion_probability)
                         for ticket_id in chosen)
    return selection

records = [(f"N-{label}-{index}", "North", label)
           for label in (0, 1) for index in range(7)]
records += [(f"S-{label}-{index}", "South", label)
            for label in (0, 1) for index in range(9)]
sample = draw_validation_ids(records, per_cell=3, seed=47)
assert len(sample) == 12
assert {probability for _, cell, probability in sample
        if cell[0] == "North"} == {3 / 7}

Performance and operating cost

For N frame records, cell construction costs O(N) expected time and O(N) space; sorting IDs for reproducible draws costs O(N log N) total in the worst case. Human adjudication and missing audit trails dominate operating cost.

Common Mistakes

  • Do not treat an oversampled validation set as if it were a simple random sample.
  • Do not reveal the desired policy result to adjudicators when it can be withheld.
  • Do not silently remove sampled records that cannot be reviewed.

Read next

Continue the workflow: PU class prior and confirmation-rate sensitivity.

ai-data
data-science
Storage details