Skip to content

Production Incident

This playbook defines how the team responds to production incidents.

SeverityMeaningExample
SEV1Major production outage or data risk.Client workflow unusable, data corruption risk, auth outage.
SEV2Significant degradation with workaround.Core workflow slow or partially broken.
SEV3Limited impact or non-critical issue.Isolated bug, minor integration failure.
RoleResponsibility
Incident CommanderCoordinates response and decisions. Usually Project Technical DRI or available senior engineer.
Technical ResponderInvestigates and fixes the issue.
Communications OwnerHandles client-facing updates. Usually Deployment Strategist.
Timeline OwnerRecords timestamps, actions, and decisions. Can be the Incident Commander on small incidents.

On small teams, one person can hold multiple roles, but the Incident Commander and Communications Owner should be explicit.

  1. Declare incident and severity.
  2. Assign roles.
  3. Create or update the Notion Incident Log entry.
  4. Confirm client impact.
  5. Stabilize the system.
  6. Communicate status internally.
  7. Communicate to client if client impact exists.
  8. Resolve or mitigate.
  9. Confirm recovery.
  10. Create follow-up tasks.

Record:

  • Detection time.
  • Impact start time, if known.
  • First response time.
  • Key investigation findings.
  • Mitigation time.
  • Recovery time.
  • Client updates sent.
  • Follow-up owners.

Internal updates should include:

  • Current impact.
  • Current hypothesis.
  • Owner.
  • Next update time.

Client updates should be factual and concise:

  • What is affected.
  • What Calibrax is doing.
  • Whether there is a workaround.
  • When the next update will be sent.

Do not speculate in client updates.