The first time production breaks badly, your team will improvise. Someone will notice. Someone will Slack about it. People will jump on a call. Maybe a fix ships. Maybe customers find out before you tell them. Eventually it ends.
Improvisation works for a while. It stops working around the time you have three engineers, paying customers, and an SLA. After that point, an incident management process is the difference between 90-minute recoveries and 6-hour ones. If you want the process stood up and someone on call who has run incidents before, that is an incident response retainer.
The components you actually need
A workable incident process at startup scale has five pieces. Not 50.
1. Detection. Alerts go to a single channel that a human is responsible for at any given time. On-call rotation with PagerDuty, BetterStack, or Opsgenie. The on-call engineer is the first responder for every alert, no matter the severity.
2. Declaration. When the on-call engineer determines this is a real incident (vs a flaky alert), they declare it. A simple Slack command (/incident in your incident management tool) creates a dedicated channel, pings the right people, and starts the timeline.
3. Coordination. One person is the Incident Commander. Their job is to coordinate, not to debug. They run the bridge, decide on actions, assign owners, communicate to stakeholders. Often this is the on-call engineer for small incidents. For severe incidents, it should be a more senior person.
4. Communication. Customer communication on the status page within 15 minutes of declaring an incident, even if you do not know much yet. Internal updates in the incident channel every 30 minutes during active incidents. Customer follow-up email after resolution.
5. Postmortem. Within five business days of resolution. Blameless, focused on systems and decisions, with concrete action items and owners. Reviewed in a 30-minute meeting with the team.
Severity levels that make sense
Three levels is enough at startup scale. More than that and you spend energy debating severity instead of fixing the problem.
- P1: Customer-facing outage or data risk. Wake people up. All-hands until resolved. Customer comms required.
- P2: Significant degradation, but customers can still mostly work. Wake people up during business hours. Coordinated response. Customer comms required.
- P3: Minor degradation, internal issue with customer-facing impact pending. Business hours only. Single owner. Update status if customer-visible.
The roles, even on a small team
Even with only 4 engineers, separate the roles during an incident.
Incident Commander: coordinates, does not debug. The brain.
Subject Matter Expert: actually debugging the issue. The hands.
Communicator: writes the status page updates, emails customers, fields stakeholder questions. The voice.
Scribe: keeps the timeline. Notes every action, every observation, every decision. The memory.
For a small incident, one person can wear multiple hats. For a serious one, the roles split. The Commander never debugs because the cognitive load of running the bridge and writing code at the same time is what causes 90-minute incidents to stretch to 6 hours.
What the postmortem must do
Blameless postmortems are not "no one is at fault." They are "the goal is to learn, not to assign blame." Specific decisions and actions are discussed in detail. Systems failures are surfaced. Individuals are not named in the postmortem document.
Three sections that matter:
Timeline. Minute by minute, what happened. What was observed, what was done, what worked, what did not.
Contributing factors. Systems, processes, and decisions that contributed to the incident. Plural. Always plural; there is never a single cause.
Action items. Specific, owned, dated. Tracked to completion in the same place you track product work. The number of postmortems with action items that never close is the leading indicator of repeat incidents.
The cultural piece
The hardest part of incident management is not the runbook. It is building a culture where incidents are normal and learning-focused rather than shameful and hidden. Teams that punish incidents have fewer reported incidents and more catastrophic ones. Teams that treat incidents as the cost of running a real system get better at running real systems.
Building an incident process?
We help startups stand up incident management, on-call rotations, and postmortem culture. Two-week engagement, working with your existing tools.
Get startedSecurity Incidents Do Not Follow the Availability Playbook
Most incident processes are written by people thinking about outages, and then get used for the first time on a suspected compromise, where several of the instincts are wrong. In an outage you restore service as fast as possible. In a security incident, the fastest restoration can destroy the evidence you need to work out what happened and what you are legally required to tell people.
Three differences matter enough to write into the process before you need them. First, containment before remediation: isolate the affected host or revoke the credential, but preserve the instance, the logs and the memory state rather than terminating and replacing. Second, a smaller room: an availability incident channel with fifteen people in it is fine, and a suspected insider incident with fifteen people in it is a problem. Have a private channel pattern and a rule for who opens it. Third, the clock is legal, not operational. Under PIPEDA you must report a breach of security safeguards to the Privacy Commissioner as soon as feasible where it creates a real risk of significant harm, and keep records of every breach whether or not it met that bar. Quebec's Law 25 has its own notification duty and its own register. Contractual clocks are often tighter than the statutory ones, and enterprise customers routinely negotiate 24 or 48 hour notification into the security addendum.
The practical consequence is that your declaration step needs a branch. When the on-call engineer declares, they pick availability or security. Security pulls in whoever owns legal and privacy, starts the evidence preservation checklist, and starts a decision log about notification that you will be very glad to have three months later.
Alert Quality Is the Process
Teams spend weeks on severity definitions and no time on the thing that determines whether the process works, which is whether the pages are real. If the on-call engineer is woken four times a week and three of those are noise, they will start acknowledging pages half asleep without reading them, and the fourth one will be the real outage.
Track two numbers per rotation and review them in the same meeting where you review postmortems. Pages per shift, and the percentage of pages that resulted in an action. If actionable pages fall below roughly half, stop building new alerts and start deleting them. An alert that nobody can act on at 3 AM belongs on a dashboard, not on a pager.
The related discipline is threshold honesty. Alerts wired to a symptom customers feel, such as checkout error rate, hold their value. Alerts wired to a cause, such as CPU above 80 percent, drift into noise because the same number means something different after every scaling change. Symptom alerts also produce better declarations, since the engineer starts with a statement of customer impact rather than a metric to interpret.
What Auditors and Enterprise Buyers Actually Test
Incident management shows up in SOC 2 under the criteria covering identification, evaluation and communication of security events, and the testing is more mundane than teams expect. An auditor asks for the policy, then asks for a list of incidents in the period, then samples two or three and asks you to walk them through from detection to closure with the artifacts to match.
That sampling exposes a specific failure. Companies keep a beautiful policy and no incident register, because their incidents lived in Slack channels that got archived. Keep a register: date, severity, short description, detection method, resolution time, whether customers were notified, and a link to the postmortem. Ten rows in a spreadsheet is enough at startup scale and it turns a difficult fieldwork conversation into a five minute one. A register that says zero incidents for the entire period, incidentally, reads as a detection failure rather than a clean year, and experienced auditors treat it that way.
Enterprise buyers ask a slightly different set of questions during vendor review. How quickly will you tell us, in hours, and does that clock start at detection or at confirmation. Who is the named contact and is it a person or an alias. Do you run tabletop exercises and when was the last one. Will you give us the postmortem for an incident that affected our data. Decide your answers in advance, because deciding them in the middle of a questionnaire tends to produce commitments your process cannot actually meet. If you promise notification within 24 hours of detection, you have just made your detection capability a contractual obligation.
The Tabletop That Is Worth Running
Most tabletop exercises are theater. A facilitator reads a scenario, everyone agrees they would follow the runbook, and a report gets filed. The version worth two hours of your team's time is narrower and more uncomfortable.
Pick a scenario with an ambiguous start, because ambiguity is where real processes fail. A support ticket says one customer can see another customer's data. Nobody is sure yet whether it is a bug in a filter or a genuine authorization flaw. Run the first thirty minutes in real time: who gets paged, who declares, what severity, what is the first thing you check, do you take the feature offline while investigating, and what do you tell the customer who reported it while you still do not know.
Then break the exercise deliberately. The person who owns that service is on a flight. The logs for the relevant table have 14 day retention and the ticket refers to something from last month. The value is in finding the dependency on one person and the gap in retention, both fixable in a week and neither visible in a well-behaved rehearsal. Write the findings into your backlog with owners and dates, exactly as you would postmortem actions.
When You Should Not Hire Anyone For This
An incident management process is one of the few things a small engineering team can genuinely build itself, and we tell people that regularly.
If you have fewer than five engineers, one product, and no contractual notification commitments, write the process yourself in an afternoon. Three severity levels, a rotation, a status page, and a rule that every P1 gets a written review. Buying a two week engagement to produce that is spending money on formatting. The publicly available material on this is good and the mechanics are not proprietary.
If your real problem is that alerts are noisy or that one engineer is the only person who understands production, no process document will fix it. That is an engineering investment and an on-boarding investment, and a consultant writing a runbook on top of it produces a document that describes a system nobody else can operate.
Outside help earns its cost in three situations. When a buyer or auditor has put a date on it and you need the register, the policy and the evidence to line up under sampling. When you need someone experienced reachable at 2 AM because your team is four people and cannot sustain a real rotation, which is what an incident response retainer is for. And when the incident you are managing is a suspected compromise and you have never handled one, where the cost of a wrong first hour is measured in evidence you can no longer recover.
If you are somewhere in between, the cheapest useful thing is usually a review rather than a build. Send us your existing process and incident register and we will tell you what an auditor will pick at and what a buyer will push back on. Our pricing is published, so you can work out whether the arithmetic favors doing it in house, and the free Workspace includes somewhere to keep the register so it stops living in archived Slack channels. If you want the process reviewed against a specific framework rather than in general, that sits inside compliance work rather than an engineering engagement.
Before you need it. Incident response on retainer means the contracts, the access and the runbooks already exist when the pager goes off.
See how a retainer worksOr talk about a retainer