Release a customer analytics mart with declared field sensitivity, tested row policies, named approvals and measurable cost per published row.
Project: release a governed customer mart
Specify the product boundary
The mart has one row per customer and reporting day, with region, eligible-order count and net captured cents. Raw landing retains namespaced customer IDs under a restricted role; analysts see regional aggregates and masked tokens. Record field classifications, allowed uses, retention, consumer roles and the business owner before creating the serving view.
Build and reconcile the release
Pin source snapshots, deduplicate captures by stable ID, resolve the dimension version valid at event time and aggregate at the declared grain. Reconcile source and mart cents by day, then publish one immutable snapshot. Grain rules prevent capture multiplication. Keep quarantine counts and unresolved customer keys visible to the operator.
Test permissions as consumers
Query the actual serving view as north-region analyst, west-region analyst, finance service and an unauthorized identity. Assert exact allowed row IDs, masked token values and denied access to raw landing. Try a join and export path available to each role. Policy tests must fail when a new sensitive field appears without an explicit allowlist decision.
Approve and cost the change
Introduce a new postal-zone field and run classification review before exposing it. The business owner approves its meaning; the steward and security approver decide access. Compare old and new output on the same input and retain a rollback pointer. Attribute compute attempts, storage generations and validation scans to this release, then report cost per accepted million rows.
Deliver audit evidence
Provide the source contract, data classification matrix, role-query transcript, field lineage, signed change record, reconciliation report, cost attribution and snapshot manifest. Simulate a denied analyst query and a failed compute attempt; neither should leak raw data or disappear from the cost ledger. Record who can authorize a repair or rollback.
Implementation
captures = [
{"id": "cap-47", "customer": "cust-7", "region": "north", "cents": 5200},
{"id": "cap-48", "customer": "cust-8", "region": "west", "cents": 1700},
]
def regional_mart(rows):
identities = set()
totals = {}
for row in rows:
if row["id"] in identities:
raise ValueError("duplicate capture")
identities.add(row["id"])
totals[row["region"]] = totals.get(row["region"], 0) + row["cents"]
return totals
assert regional_mart(captures) == {"north": 5200, "west": 1700}Performance and operating cost
The reference aggregation takes O(N) expected time and O(N + G) memory for N captures and G regions because it checks identities. Production cost includes versioned joins, access-policy evaluation, release validation, retained snapshots and failed attempts. Report cost against accepted output and keep policy tests in the release gate even when their scans add expense.
Common Mistakes
- Do not publish a new derived identifier before classifying it.
- Do not validate permissions with an administrator account.
- Do not price only the final successful compute attempt.
Read next
- Data classification and access boundaries
- Row policy and column mask tests
- Dataset ownership and change approval
- Pipeline cost attribution and right-sizing
- Project: release a reconciled revenue mart
Continue the workflow: Serving indexes and freshness contracts.
Continue the workflow: Project: prove erasure after a regional restore.
Continue the workflow: Project: protect analytics reads during a tenant burst.
Continue the workflow: Project: release regional metrics without exposing small groups.
Continue the workflow: Project: audit a payment column change across a mart.
