Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Small-group suppression and release grain

Last updated: 6 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

An aggregate can expose a person when its released group is too narrow or when another published total lets a reader subtract a hidden group.

Name the privacy unit

A customer may make several transactions. Counting rows instead of distinct customers makes a small group look safer than it is. Define the privacy unit and the smallest combination of dimensions that a consumer may query. A region-day-channel cube is far more revealing than a region-month rollup, even when both report the same overall revenue. Fact grain must be separate from the release grain.

Set a minimum cohort rule

Require a minimum distinct-person count for each released cell and suppress cells below it. State whether totals that include suppressed cells can also be released. A single suppressed cell can be recovered by subtracting visible child cells from an exact parent total. Test the full set of available queries and exports, not one dashboard tile in isolation. A threshold reduces risk but does not prove anonymity.

Generalize when suppression destroys use

If store-day cells are sparse, publish store-week or region-day groups instead. Bucketing can preserve useful trends while making individuals harder to isolate. Keep the grouping decision and owner approval in the dataset contract. Never choose a broader group by silently merging unrelated populations if that changes the business meaning of the metric.

Control repeated releases

A cell with 47 people today and 48 tomorrow may disclose the added person when exact totals are compared. Define a release cadence and whether historical revisions replace earlier public answers. For high-risk use, consider a formally reviewed privacy mechanism with a budget for repeated queries; simple thresholding cannot solve every differencing attack. Correction handling still has to keep the published numbers accurate.

Test the consumer surface

Query permitted groups, suppressed groups, parent totals and neighboring dates through the same API the consumer uses. Inspect CSV exports, cached responses and chart tooltips. A privacy check only in the transformation job is bypassed if a flexible query endpoint can create narrower groups. Record the accepted release schema and deny unapproved dimensions before any result leaves the service.

Implementation

python
customers_by_region = {
    "west": {"cust-47", "cust-48", "cust-49", "cust-50", "cust-51", "cust-52"},
    "east": {"cust-53", "cust-54"},
}

def releasable_counts(groups, minimum_people):
    return {region: len(customers) for region, customers in groups.items()
            if len(customers) >= minimum_people}

assert releasable_counts(customers_by_region, 5) == {"west": 6}

Performance and operating cost

Counting distinct units across N records is O(N) expected time and O(U) state for U distinct people in the reference. Fine-grained cubes multiply stored groups and disclosure checks; broader groups cost less but lose detail. This simple gate is illustrative, not a formal privacy guarantee: repeated releases and overlapping totals require additional policy and testing.

Common Mistakes

  • Do not count transactions when the privacy unit is a person.
  • Do not publish an exact parent total that reveals one suppressed child by subtraction.
  • Do not claim a group threshold alone guarantees anonymity.

Read next

ai-data
data-engineering
Storage details