A privacy review should test what an adversary can infer from the released interface and state what the test does not cover.
Model leakage evaluation: membership risk, memorization and limits
Set attacker access
A public scoring endpoint, a downloadable model and an internal embedding index expose different information. Specify query limits, output granularity and whether the attacker knows a learner’s candidate record. A confidence score with many decimal places can reveal more than a coarse decision. The threat model determines which attack simulation is relevant.
Use held-out comparisons
A simple membership probe compares model behavior on training and genuinely held-out records matched by time and population. Large performance gaps can be a warning of memorization, but they do not prove a particular record leaked. Conversely, a weak probe does not prove safety against stronger attacks. Keep the audit data separate from the model-selection loop.
Examine rare records
Unusual receipt text, uncommon names or a learner’s distinctive sequence may be reproduced by a generative model or close neighbor search. Test canary-like synthetic strings under a controlled procedure, with no real sensitive text in public reports. Review logging, caches and retrieval indexes as well as model weights; leakage can come from surrounding systems.
Report defensible findings
Compare attack success with a population-aware baseline, provide confidence intervals and state query budget. If a membership probe reaches 58 percent accuracy against a balanced 50 percent baseline, report the sample size and uncertainty rather than declaring a universal breach. Record what interface was tested and what remains untested before a release decision.
Implementation
def membership_probe_accuracy(predicted_membership, actual_membership):
if len(predicted_membership) != len(actual_membership) or not actual_membership:
return None
correct = sum(predicted == actual for predicted, actual
in zip(predicted_membership, actual_membership))
return correct / len(actual_membership)Performance and operating cost
Scoring N probe cases costs O(N) time and O(1) auxiliary space. A real attack evaluation may require many model queries and strict test access controls; its result is evidence about the tested interface, not a proof of overall privacy.
Common Mistakes
- Do not interpret one weak attack as proof of safety.
- Do not publish real sensitive canaries in an audit report.
- Do not ignore retrieval indexes and logs when reviewing model leakage.
