A service catalog is an indexed description of software and its relationships. Its owner field is a claim, not a pager test. A renamed team, deleted repository, stale on-call schedule, or missing runbook can leave a polished catalog entry that fails during an incident. Treat catalog metadata as a contract whose critical fields must be checked against systems of record and against an actual escalation path. The service identifier should remain stable across repository renames so incidents and release evidence still join to the same workload.
Service catalog ownership: prove that the named team can respond
Operational decision
For a payment-reconciliation API, register a stable service ID, repository, production workload, owning team, escalation schedule, runbook, data classification, and retirement state. A scheduled verifier resolves the team and schedule through their current identities and records the last successful check; it does not infer ownership from the person who last committed code. Inject three faults into a disposable catalog: delete the schedule, move the repository, and disable the runbook link. Each should flag a different broken contract with an assigned repair owner. During a simulated incident, start from a runtime alert and reach the responsible human without searching chat history. When ownership transfers, the receiving team acknowledges the alert path before the old team is removed, and the record retains a transfer date. A stale entry is marked suspect even if its source file still parses.
Service: payment-reconciliation
Stable ID: payments.reconciliation
Owner: team-ledger-operations
Escalation: current rota ID and tested contact
Runbook: current incident procedure
Runtime: production workload identity
Last verified: timestamp and result
Transfer: receiving-team acceptance requiredCost and verification
Verifying S services against external systems takes O(S) lookups per field, before caching or batching. Aggressive full-catalog polling can overload identity and paging APIs; use incremental checks with a periodic full audit. Measure unresolved owners, failed escalation drills, stale verification age, and incidents where ownership lookup exceeded the response target. A catalog entry is not evidence of response readiness until the lookup path works.
Common Mistakes
- Do not use the last code author as the production owner.
- Do not accept a team name whose paging destination no longer exists.
- Do not remove the previous owner before the new team accepts escalation.
Connected lessons
- DevOps: delivery, infrastructure, and reliable operations
- Alert design: page on impact and include a first action
- Runbook automation: put a stop gate before the irreversible step
- Release evidence: tie one deployed digest to one approval decision
- Cloud cost and capacity: assign an owner to each recurring resource
