Production incident escalation
You are a production systems incident response specialist. Help me structure/manage: [CONTEXT — an incident is happening NOW in [SYSTEM] and I need structure to handle it well/I want to set up the team's incident response process before it happens again/we had a poorly managed incident recently and want to fix the process]. Deliver — if incident is active now: the first steps in the right order (formally declare the incident with estimated severity — delay in declaring is delay in mobilizing help; name an incident commander — the person coordinating and deciding, not necessarily the one fixing it technically, so whoever is debugging doesn't also have to manage communication; open a single incident communication channel — avoid information fragmented across DMs and parallel threads; stabilize before investigating deeply — immediate mitigation as priority over understanding root cause completely); for the structured process: severity levels defined with objective criteria (what makes an incident SEV1 versus SEV3 in MY context — user impact, scope, reversibility — with the response protocol corresponding to each level, including who gets called and with what urgency), clear roles during an incident (incident commander coordinating, technical team executing, external/internal communication as separate explicit responsibility — nobody accumulating all three roles in a serious incident), escalation with criteria (when to call more people, when to wake someone at midnight, escalation chain defined in advance to not waste time deciding this IN the middle of crisis), communication during the incident (status updated at regular intervals even without news — 'still investigating, next update in 15 min' is better than silence; external customer communication when applicable, honest without unnecessary panic), formal incident closure (clear criterion for when to consider it resolved — not just 'seems normal' without confirmation), blameless post-mortem as a mandatory and non-optional stage (timeline reconstructed, root cause identified through systematic investigation, preventive actions with owner and deadline — and a blameless culture that makes people report problems early instead of hiding from fear), and the documented and tested incident runbook before the next crisis (the living document the team reviews and simulates periodically, not written once and forgotten in a drawer). Objective: a crisis handled with calm and method — that ends up resolved quickly and becomes real learning, not a second communication crisis on top of the first one.