Issuing a new certificate with the same compromised key does not remove the attacker's ability to impersonate the service. Generate a new key under protected control, issue a matching certificate, remove exposed copies, and determine whether the issuer or only one leaf was affected. The response also examines active connections, session resumption, access logs, and dependent clients. Revocation can support containment where clients enforce it, but its actual effect must be tested.
TLS key compromise: replace authority and end stale sessions
Operational decision
A settlement API key appears in a public CI artifact. The incident lead freezes the affected artifact, identifies every environment where the key is mounted, and records the certificate fingerprint and first known exposure time. The team generates a new key in its approved secret boundary, issues a new leaf, distributes it to a canary listener, and verifies fresh handshakes from required clients. It then rolls all listeners, drains long-lived connections on a bounded schedule, and rotates session-ticket material if the termination layer and threat model require it. Inspect registry, CI logs, build caches, and backups for other copies without copying the key into incident notes. Revoke the old leaf where the PKI supports it, but also test a client class that lacks revocation checking and rely on key replacement plus trust or identity controls there. If the CA signing key itself is suspect, activate the root or intermediate rollover plan rather than merely replacing one leaf. The incident record includes which certificate is currently served per endpoint and when the old identity ceased to be accepted. Do not claim complete containment solely because the Kubernetes Secret changed.
Settlement key incident
Scope: leaked key fingerprint and artifact locations
New identity: fresh key plus issued leaf
Deployment: canary, all listeners, standby region
Sessions: old connection and ticket behavior checked
Revocation: enforcement tested by client class
Cleanup: CI, registry, cache, and backup copies reviewed
Exit: old identity rejected on required pathsCost and verification
Scanning A artifacts and E endpoints is at least O(A + E) investigative work; the real duration depends on retained logs, distribution lag, and client sessions. Emergency rotation may add handshake load and brief connection churn, so size spare capacity before draining. Measure time from exposure discovery to new key serving, old fingerprint presence, stale session count, and remaining accessible copies. Keep containment proof tied to actual endpoints, not a ticket checkbox.
Common Mistakes
- Do not reissue a new certificate over the exposed private key.
- Do not paste the compromised key into incident documentation.
- Do not close the incident at Secret update while old listeners remain active.
Connected lessons
- DevOps: delivery, infrastructure, and reliable operations
- Credential incident response: revoke access before rebuilding trust
- Certificate reloads: distinguish file delivery from active TLS state
- Certificate revocation: test the client behavior and status-service dependency
- Software supply chain: SBOM and provenance at admission
