A treatment policy converts estimated incremental benefit into an action under capacity, cost, eligibility and safety limits.
Capacity-constrained treatment policies and guardrails
Convert probability uplift into net value
If one additional course completion is worth 24 units to an outreach team and one reminder costs 0.7 units, a learner with predicted uplift 0.06 has estimated net value 24 times 0.06 minus 0.7 = 0.74 units. A learner at 0.02 yields negative 0.22 units and should not be targeted merely because capacity remains. Units must describe a real decision objective, not an arbitrary score scale.
Apply eligibility before ranking
Exclude opted-out learners, recently contacted learners and accounts with immature outcomes before computing a top-K list. The eligibility contract is versioned and shared by training, evaluation and serving. A model that was evaluated on one eligible population but served on another has no directly transferable policy-value estimate. The event contract defines this population.
Respect budget and tie rules
With a limit of 24 reminders per day, select positive-net-value learners until capacity is used. For deterministic replay, sort ties by a stable learner ID after net value. If the outreach team imposes per-region ceilings, solve the constrained selection explicitly and evaluate that exact rule, not an unconstrained top-24 proxy. Offline policy value must use the deployed policy.
Keep a measurement lane
Retain a randomized control group among eligible learners so that changes in lesson difficulty, exam season or message fatigue can be separated from model drift. Do not treat the control lane as leftover capacity. Monitor opt-outs, complaint rate, delivery failures and group-level effects alongside completions. A high average uplift does not excuse a severe harm in a protected or operationally important slice.
Define rollback and refresh
Freeze the model, feature snapshot, value conversion, capacity and policy version in each decision log. Pause targeting if assignment logs fail, overlap falls below the release threshold, or a guardrail breaches its limit. Update the model only after outcome windows mature and a fresh randomized evaluation supports the change. The release project gives a concrete decision packet.
Implementation
eligible_learners = [
{"learner": "acct-47", "uplift": 0.06, "opted_out": False},
{"learner": "acct-62", "uplift": 0.02, "opted_out": False},
{"learner": "acct-83", "uplift": 0.11, "opted_out": True},
{"learner": "acct-94", "uplift": 0.04, "opted_out": False},
]
def choose_reminders(records, capacity, value_per_completion=24, reminder_cost=0.7):
if capacity < 0:
raise ValueError("capacity must be nonnegative")
candidates = []
for record in records:
net_value = value_per_completion * record["uplift"] - reminder_cost
if not record["opted_out"] and net_value > 0:
candidates.append((net_value, record["learner"]))
candidates.sort(key=lambda candidate: (-candidate[0], candidate[1]))
return [learner for _, learner in candidates[:capacity]]
assert choose_reminders(eligible_learners, 2) == ["acct-47", "acct-94"]Performance and operating cost
For N eligible learners, sorting costs O(N log N) time and O(N) space; a bounded heap can reduce selection to O(N log K). Constrained optimization, control-lane allocation and monitoring add operational cost. The largest risk is an overconfident uplift estimate near the net-value threshold.
Common Mistakes
- Do not fill capacity with negative-net-value actions.
- Do not evaluate one eligibility rule and serve another.
- Do not remove the randomized control lane after an early positive result.
