A membership exposure audit asks whether model outputs help distinguish training records from comparable nontraining records.
Membership exposure audits: test train-versus-holdout distinguishability
Define a narrow question
A retail-return classifier predicts whether a return needs manual review. The privacy team wants to know whether its exposed probabilities reveal something about which cases were used for training. Build a permitted audit with known member and nonmember records, then fix the observable output fields and a predeclared attack statistic. The result describes this audit frame and access level, not every possible adversary. API output scope determines what an ordinary caller can observe.
Make cohorts comparable
Members and nonmembers should represent the same acquisition period, return types and labeling procedure where possible. A naive comparison of old training returns against newly collected returns can detect time drift rather than membership. Group related transactions by account to avoid near duplicates across cohorts. Keep audit samples separate from model tuning and threshold selection. If the attacker has only route labels, audit that surface; do not score with internal logits and present the result as public API risk. Split replay helps preserve cohort identities.
Predeclare a test and report uncertainty
One simple diagnostic marks an example as a likely member if its exposed confidence exceeds a fixed threshold. Count true-positive and false-positive rates on balanced, matched cohorts, and report their difference with sample sizes. A positive gap merits investigation; a zero gap does not prove privacy. More capable attacks or subgroup effects may still exist. Repeat across model revisions and output scopes without tuning the threshold on the same protected audit results. Mitigation review compares the same evaluation plan after a change.
Handle audit data carefully
Audit records themselves can contain sensitive return details. Use approved identifiers, restrict raw inputs, separate the audit ledger from public telemetry and set retention. Store model digest, output scope, cohort definition, test threshold and result ID. Do not release record-level membership guesses to general dashboards. Inference logging limits apply here too; the project demonstrates why unmatched cohorts make a dramatic but misleading result.
Implementation
def fixed_threshold_membership_audit(members, nonmembers, threshold):
if not members or not nonmembers:
return {"state": "hold:empty-cohort"}
member_hits = sum(score >= threshold for score in members)
nonmember_hits = sum(score >= threshold for score in nonmembers)
true_positive_rate = member_hits / len(members)
false_positive_rate = nonmember_hits / len(nonmembers)
return {"state": "measured", "member_count": len(members),
"nonmember_count": len(nonmembers),
"true_positive_rate": true_positive_rate,
"false_positive_rate": false_positive_rate,
"rate_gap": true_positive_rate - false_positive_rate}
report = fixed_threshold_membership_audit(
[0.82, 0.74, 0.43, 0.91], [0.39, 0.58, 0.72, 0.31], 0.70)
assert report["true_positive_rate"] == 0.75
assert report["false_positive_rate"] == 0.25
assert report["rate_gap"] == 0.5
Performance and operating cost
This fixed-threshold diagnostic is O(m + n) time and O(1) extra space for m members and n nonmembers. A fuller audit costs additional model queries, matched labeling and controlled review. Small cohorts produce noisy rates; a single low-cost test is useful for triage but cannot certify privacy.
Common Mistakes
- Comparing old training data with a newer holdout and calling time drift membership leakage.
- Selecting the threshold after seeing the protected audit result.
- Testing internal model logits while claiming risk from a route-only public API.
- Treating a low measured gap as proof that no privacy attack can succeed.
Read next
- Privacy release gates: reduce exposed detail and retest model utility
- Project: audit privacy exposure in a return-review model
- Prediction API exposure: define what a client may learn from scores
- Inference logs: keep diagnostic joins without copying sensitive payloads
- Training replay: freeze the cohort, split and runtime
