Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Privacy-aware ML: threat model, data minimization and purpose

Last updated: 5 Oct 20265 min read
tutorial
IntermediateBy AITrove Editorial

A privacy claim needs a defined adversary, protected unit, data flow and retention boundary before a technique can be evaluated.

Name the protected unit

For an AI Trove lesson model, protecting a learner may mean preventing an observer from inferring that learner’s activity from a released model. That differs from preventing an internal operator from viewing raw events or another tenant from receiving a feature. Write the adversary’s access: public model outputs, logs, shared gradients or database queries. Inference logging is one part of the exposure path.

Reduce collection first

List the fields required for the decision and avoid copying unrelated profile details into training. Replace raw text with a task-specific signal where the loss is acceptable; cap retention and keep a deletion key. A hashed learner ID is still a persistent linkable identifier, not automatic anonymity. Separate collection permission, training permission and publication of aggregate results.

Trace derivatives

A raw lesson event may generate a warehouse row, feature vector, embedding, training set, model checkpoint and diagnostic log. Inventory those descendants before promising deletion or restricted reuse. Feature lineage and model manifests should identify the version and purpose of each derivative without placing sensitive values in the manifest itself.

Test a policy boundary

Create one learner who permits feed personalization but opts out of model training. Their request-time state may serve the allowed feed while their events are absent from the training extract. Add a withdrawn learner and verify the online feature key, cached vector and future batch inputs are excluded under the defined policy. A reviewer should see which artifacts remain and why.

Implementation

python
def training_eligible(events):
    return [event for event in events
            if event["training_permission"] and not event["withdrawn"]
            and event["purpose"] == "lesson-recommendation"]

Performance and operating cost

Filtering N events is O(N) time and O(A) output memory for A allowed events. The real cost is maintaining permission and deletion state consistently across derived datasets; a row filter alone is not a privacy guarantee.

Common Mistakes

  • Do not call a hashed stable ID anonymous.
  • Do not reuse data for training merely because it was collected for serving.
  • Do not promise deletion before inventorying derived artifacts.

Read next

Continue the workflow: Synthetic-data privacy: copies, nearest neighbors and disclosure.

ai-data
privacy-aware-ml
Storage details