A privacy claim needs a defined adversary, protected unit, data flow and retention boundary before a technique can be evaluated.
Privacy-aware ML: threat model, data minimization and purpose
Name the protected unit
For an AI Trove lesson model, protecting a learner may mean preventing an observer from inferring that learner’s activity from a released model. That differs from preventing an internal operator from viewing raw events or another tenant from receiving a feature. Write the adversary’s access: public model outputs, logs, shared gradients or database queries. Inference logging is one part of the exposure path.
Reduce collection first
List the fields required for the decision and avoid copying unrelated profile details into training. Replace raw text with a task-specific signal where the loss is acceptable; cap retention and keep a deletion key. A hashed learner ID is still a persistent linkable identifier, not automatic anonymity. Separate collection permission, training permission and publication of aggregate results.
Trace derivatives
A raw lesson event may generate a warehouse row, feature vector, embedding, training set, model checkpoint and diagnostic log. Inventory those descendants before promising deletion or restricted reuse. Feature lineage and model manifests should identify the version and purpose of each derivative without placing sensitive values in the manifest itself.
Test a policy boundary
Create one learner who permits feed personalization but opts out of model training. Their request-time state may serve the allowed feed while their events are absent from the training extract. Add a withdrawn learner and verify the online feature key, cached vector and future batch inputs are excluded under the defined policy. A reviewer should see which artifacts remain and why.
Implementation
def training_eligible(events):
return [event for event in events
if event["training_permission"] and not event["withdrawn"]
and event["purpose"] == "lesson-recommendation"]Performance and operating cost
Filtering N events is O(N) time and O(A) output memory for A allowed events. The real cost is maintaining permission and deletion state consistently across derived datasets; a row filter alone is not a privacy guarantee.
Common Mistakes
- Do not call a hashed stable ID anonymous.
- Do not reuse data for training merely because it was collected for serving.
- Do not promise deletion before inventorying derived artifacts.
Read next
- Privacy units and bounded contribution: count people, not events
- Privacy operations: deletion lineage, access gates and release review
- Inference logs: keep diagnostic joins without copying sensitive payloads
- Feature definitions: versioned views, ownership and lineage
Continue the workflow: Synthetic-data privacy: copies, nearest neighbors and disclosure.
