A dataset owner accepts its meaning, access and change process; a pipeline maintainer owns the job that computes it, and consumers own their use of it.
Dataset ownership and change approval
Name the decision makers
For a customer mart, record a business owner for field meaning, a steward for classification, a technical maintainer for pipeline health and a security approver for new sensitive access. One person can fill several roles in a small team, but each decision still needs an accountable name. An unowned table should not silently become a production dependency.
Make the contract reviewable
Store row grain, keys, update cadence, freshness target, null meaning, correction policy, allowed uses and retention beside the dataset. Link the upstream sources and known consumers through lineage. A consumer should be able to determine whether a changed field is safe without searching chat history or reverse engineering a dashboard query.
Classify changes by effect
A new optional field may be backward-compatible for readers but still change classification or cost. Renaming a key, altering its grain or changing a metric denominator is a breaking semantic change even if the schema remains valid. Schema rollout handles code compatibility; owner approval covers business meaning and access.
Stage a measurable migration
Publish the proposed contract and sample output, notify known consumers, run old and new calculations on the same pinned input, and record count or metric differences. Set a dual-read or deprecation window when needed. Promotion requires named approval, validation evidence and a rollback pointer. A pass from automated tests alone cannot approve a new use of sensitive data.
Keep evidence after release
Attach change ID, approving roles, input snapshot, code revision, access-policy revision and consumer communication to the release manifest. If a downstream report changes, the operator can explain whether the cause was new source data, a model change or a policy change. Replay manifests provide the technical half of that explanation.
Implementation
required_approvals = {
"grain_change": {"business_owner", "technical_maintainer"},
"new_sensitive_field": {"data_steward", "security_approver"},
}
def release_approved(change_type, signed_roles):
required = required_approvals[change_type]
return required.issubset(set(signed_roles))
assert not release_approved("new_sensitive_field", ["data_steward"])
assert release_approved("new_sensitive_field", ["data_steward", "security_approver"])Performance and operating cost
An approval membership check is O(A) for A signed roles and needs little memory; the significant cost is the dual-run, consumer testing and retained release evidence. That cost should be proportional to impact. Automate discovery and comparison so humans review the meaning and authorized use, rather than manually repeating deterministic checks.
Common Mistakes
- Do not equate schema compatibility with semantic compatibility.
- Do not let a maintainer approve a new sensitive use by default.
- Do not discard the contract version once a migration finishes.
