Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Project: release a governed customer mart

Last updated: 6 Oct 20265 min read
project
AdvancedBy AITrove Editorial

Release a customer analytics mart with declared field sensitivity, tested row policies, named approvals and measurable cost per published row.

Specify the product boundary

The mart has one row per customer and reporting day, with region, eligible-order count and net captured cents. Raw landing retains namespaced customer IDs under a restricted role; analysts see regional aggregates and masked tokens. Record field classifications, allowed uses, retention, consumer roles and the business owner before creating the serving view.

Build and reconcile the release

Pin source snapshots, deduplicate captures by stable ID, resolve the dimension version valid at event time and aggregate at the declared grain. Reconcile source and mart cents by day, then publish one immutable snapshot. Grain rules prevent capture multiplication. Keep quarantine counts and unresolved customer keys visible to the operator.

Test permissions as consumers

Query the actual serving view as north-region analyst, west-region analyst, finance service and an unauthorized identity. Assert exact allowed row IDs, masked token values and denied access to raw landing. Try a join and export path available to each role. Policy tests must fail when a new sensitive field appears without an explicit allowlist decision.

Approve and cost the change

Introduce a new postal-zone field and run classification review before exposing it. The business owner approves its meaning; the steward and security approver decide access. Compare old and new output on the same input and retain a rollback pointer. Attribute compute attempts, storage generations and validation scans to this release, then report cost per accepted million rows.

Deliver audit evidence

Provide the source contract, data classification matrix, role-query transcript, field lineage, signed change record, reconciliation report, cost attribution and snapshot manifest. Simulate a denied analyst query and a failed compute attempt; neither should leak raw data or disappear from the cost ledger. Record who can authorize a repair or rollback.

Implementation

python
captures = [
    {"id": "cap-47", "customer": "cust-7", "region": "north", "cents": 5200},
    {"id": "cap-48", "customer": "cust-8", "region": "west", "cents": 1700},
]

def regional_mart(rows):
    identities = set()
    totals = {}
    for row in rows:
        if row["id"] in identities:
            raise ValueError("duplicate capture")
        identities.add(row["id"])
        totals[row["region"]] = totals.get(row["region"], 0) + row["cents"]
    return totals

assert regional_mart(captures) == {"north": 5200, "west": 1700}

Performance and operating cost

The reference aggregation takes O(N) expected time and O(N + G) memory for N captures and G regions because it checks identities. Production cost includes versioned joins, access-policy evaluation, release validation, retained snapshots and failed attempts. Report cost against accepted output and keep policy tests in the release gate even when their scans add expense.

Common Mistakes

  • Do not publish a new derived identifier before classifying it.
  • Do not validate permissions with an administrator account.
  • Do not price only the final successful compute attempt.

Read next

Continue the workflow: Serving indexes and freshness contracts.

Continue the workflow: Project: prove erasure after a regional restore.

Continue the workflow: Project: protect analytics reads during a tenant burst.

Continue the workflow: Project: release regional metrics without exposing small groups.

Continue the workflow: Project: audit a payment column change across a mart.

ai-data
data-engineering
Storage details