Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Project: release regional metrics without exposing small groups

Last updated: 6 Oct 20265 min read
project
AdvancedBy AITrove Editorial

Build a regional payment summary whose query surface, exports and lineage catalog agree on the same disclosure boundary.

Write the release contract

Create payment records for several regions and days, including a sparse region with two people making multiple payments. Define the person as the privacy unit and name the permitted region-week output grain. List every consumer: dashboard, CSV export, query API and catalog. The group gate must apply to all of them, not only the chart.

Build and challenge the aggregate

Count distinct people and sum payment amounts for each candidate cell. Set a minimum of five people, suppress the sparse region and test whether any exact parent total reveals its amount through subtraction. If it does, withhold or coarsen that parent as well. Introduce a late correction and rerun the release checks before moving the pointer. A corrected table does not authorize a narrower disclosure.

Split broad and restricted lineage

Publish dataset dependencies, release generation and opaque run ID in the general catalog. Keep partition paths, tenant mapping and detailed source positions in a restricted manifest. Metadata minimization should be visible in a failed-job log and support export too. Grant one incident-responder role temporary access and prove a general analyst is denied the restricted view.

Test differencing and replay

Publish two adjacent candidate weeks, then see whether subtracting their exact totals isolates a newly added person. Reject or generalize a release that leaks through this path. Replay the same correction and verify the published aggregate does not double count it. Query both the serving API and a cached response after pointer movement; an old cache can disclose an earlier, less protected generation.

Submit a proof bundle

Provide the privacy-unit definition, approved dimensions, suppressed-cell report, parent-total test, late-correction result, public and restricted lineage projections, role checks, cache result and final consumer query. Show one denied narrow query and one permitted broad query. The exercise is not complete merely because a transformation produced a table with fewer rows.

Implementation

python
payments = [
    ("west", "cust-47", 2375), ("west", "cust-48", 6400),
    ("west", "cust-49", 1250), ("west", "cust-50", 3100),
    ("west", "cust-51", 800), ("east", "cust-52", 900),
    ("east", "cust-52", 400), ("east", "cust-53", 1700),
]

def released_totals(records, minimum_people):
    groups = {}
    for region, customer_id, cents in records:
        state = groups.setdefault(region, {"people": set(), "cents": 0})
        state["people"].add(customer_id)
        state["cents"] += cents
    return {region: state["cents"] for region, state in groups.items()
            if len(state["people"]) >= minimum_people}

assert released_totals(payments, 5) == {"west": 13925}

Performance and operating cost

Aggregation is O(N) expected time and O(U + G) state for N payments, U unique people and G groups. Full disclosure testing adds work across all released combinations and versions, not just the current table. A restricted evidence store and dated last-good release add storage, while coarse cells reduce analytical detail. The threshold demonstration is not a formal anonymity guarantee.

Common Mistakes

  • Do not count the same person twice because they made multiple payments.
  • Do not expose a suppressed amount through a parent total or stale export.
  • Do not put tenant paths into a broadly readable lineage catalog.

Read next

ai-data
data-engineering
Storage details