Security

Real offensive depth

Testing and defence led by a published security researcher with five CVEs, including a CVSS 9.1 Mirai botnet kill-switch.

All security →
Compliance

Audit-ready, fixed scope

SOC 2, ISO, and the Canadian privacy stack, run end to end with an independent auditor.

All frameworks →
Resources

Learn the space

Original research, free tools, and plain-language guides on security and compliance, from a published security researcher.

Read the blog →
Security

How to Get Incident Response: A Step-by-Step Guide

When a breach hits, the first hour decides how bad the next month will be. Most companies find that out the hard way, scrambling to find a responder while an attacker is still inside the network. Getting incident response in place before you need it is not complicated, but it does take a deliberate process. Here is how to do it, in order, with the timelines you should actually expect.

Step 1: Figure out what you are protecting

Before you talk to anyone about incident response, know what "incident" means for your business. A SaaS company worried about customer data exposure needs different coverage than a fintech firm worried about transaction fraud or an e-commerce shop worried about payment card compromise. Write down your top three risk scenarios: ransomware, a compromised admin account, a leaked API key, whatever keeps you up at night. This list becomes the scope of your engagement.

Timeline: half a day internally, usually a conversation between the CTO and whoever owns security.

Step 2: Decide between building in-house and retaining outside help

This is the fork most companies get wrong. Building an internal security operations centre means hiring analysts, buying and tuning detection tooling, and running an on-call rotation, all before you have handled a single real incident. For most small and mid-sized companies, that is a seven-figure commitment before it pays off. An incident response retainer gets you named responders and a contracted service level agreement for a fraction of that cost, and you are covered from day one instead of eighteen months from now.

If you already have a security team and just need surge capacity for the worst-case scenario, a retainer still makes sense as backup. If you have no dedicated security staff, it is close to the only sane option.

Timeline: a few days of internal discussion, sometimes a budget approval cycle.

Step 3: Get your environment audit-ready

Responders move faster when they are not starting from zero. Before you sign anything, pull together an inventory of your systems: cloud accounts, production databases, admin access lists, and any existing logging or monitoring tools. If you already have compliance work underway, such as a SOC 2 program, a lot of this documentation already exists. If you do not, this is the point where gaps show up, missing logs, no centralized identity provider, no clear owner for a given system.

Timeline: one to two weeks, depending on how mature your existing documentation is.

Before you need it. Incident response on retainer means the contracts, the access and the runbooks already exist when the pager goes off. See how a retainer works

Step 4: Choose a provider and set the scope

When evaluating an incident response partner, ask direct questions: who exactly responds when you call, what is the guaranteed response time in the contract, and what does the retainer include versus bill separately. A named responder and a clear SLA are the two things that actually matter here. A vague promise of "24/7 support" without a contracted response window is not incident response, it is a sales pitch.

traztech's incident response retainer is built around exactly this: named responders who already know your environment, and a service level agreement in the contract, not in a slide deck. That familiarity matters more than it sounds, a responder who has never seen your architecture before spends the first hours of an incident just getting oriented.

Timeline: two to four weeks from first conversation to signed contract, including any procurement or legal review on your side.

Step 5: Onboard before anything goes wrong

A retainer is only as good as the onboarding behind it. This is where the responder gets access to your environment (or a documented path to get it fast), reviews your architecture, and agrees on escalation contacts. Good providers will also run a short tabletop exercise, walking your team through a simulated incident so everyone knows who calls whom and what happens in the first thirty minutes.

Timeline: two to three weeks for full onboarding, including at least one tabletop session.

Step 6: Keep it current

Systems change, staff turn over, and a retainer that was accurate six months ago can go stale fast. Review contact lists and system inventory quarterly, and re-run the tabletop annually or after any major infrastructure change. If your company is also pursuing certifications like SOC 2, this incident response process typically needs to be documented as part of that audit anyway, so keeping the two in sync saves duplicate work. Our compliance programs are built to work alongside an incident response retainer rather than as a separate silo.

Timeline: ongoing, roughly a quarter-day per quarter plus one annual exercise.

What this looks like end to end

Realistically, from the first internal conversation to having a fully onboarded incident response retainer in place, most companies are looking at four to eight weeks. That is far faster than the twelve to eighteen months it typically takes to stand up an internal SOC with hired analysts and tuned tooling, and it costs a fraction as much. The retainer model exists precisely because most companies do not need round-the-clock in-house staff, they need to know that when something goes wrong, a specific person picks up the phone within a contracted window.

The companies that get burned are almost always the ones that treated incident response as something to figure out later. By the time you need a responder, you do not have four weeks. You have an hour, maybe less.

Getting started

If you do not currently have a named incident response contact and a contracted SLA, that is the gap to close first, before anything else on your security roadmap. Contact traztech to talk through your environment and find out what an incident response retainer looks like for a company your size.

Check your insurance policy before you sign anything

This is the step that most often gets skipped and most often causes an expensive argument later. Many cyber insurance policies require you to use a firm from the insurer's approved panel, and to notify the insurer or their breach counsel before engaging anyone. If you call your own responder first and rack up hours, the carrier can decline to reimburse those hours, and in some policies the failure to notify promptly affects the claim itself.

So read the policy, find the incident notification clause and the panel provision, and work out which of three situations you are in. Some policies mandate the panel with no flexibility. Some allow a pre-approved alternative if you nominate them at renewal, which is a five minute conversation with your broker and worth having. Some have no panel restriction at all.

If your carrier mandates a panel firm, a retainer with an outside provider is not wasted, but it changes what you are buying. The panel firm does the incident, and your retained provider does the preparation, the runbooks, the tabletop exercises and the technical work the panel firm will not touch, such as rebuilding what was compromised. Be clear with everyone about which role they play, before the day you need both.

Reading the retainer contract properly

Retainers vary far more than the marketing suggests, and five clauses decide whether the one you signed is any good.

What triggers the SLA. A response time of one hour means nothing until you know what starts the clock and who can start it. Is it a phone number that reaches a human, or an email address monitored during business hours? Which of your people are authorised to declare an incident? If only the CTO can invoke the retainer and the CTO is on a flight, you have a coverage gap written into your own contract.

Prepaid hours and what happens to them. Many retainers include a block of hours. Ask whether unused hours expire, whether they roll over, and whether they can be spent on preparation work such as tabletops, log review or runbook writing. Hours that can only be spent during an incident are hours you are hoping to waste, and hours that expire annually reward you for having a bad year.

The rate card beyond the block. Incidents run past their estimates. Know the hourly rate for work beyond the included hours, whether out-of-hours attracts a multiplier, and whether there is a cap or an approval step before the meter runs further.

Where the work stops. Most incident response scopes cover detection, containment, investigation and a report. They frequently do not cover rebuilding systems, restoring from backups, negotiating with an extortion group, notifying regulators or customers, or the remediation project that follows. None of those exclusions is unreasonable. All of them are things you will need, so establish now who does them.

Who owns the report and who receives it. If the engagement runs through external counsel, the report may be produced under privilege, which affects who can see it and how it can be used later. If a customer or regulator will need something in writing, agree in advance that a shareable summary can be produced alongside the technical report. Retrofitting that after the fact is awkward and sometimes not possible.

The access preparation that actually saves hours

Onboarding is where the retainer earns its money, and the difference between good and poor onboarding is measured in hours of an active incident. Four items matter more than the rest.

Break-glass access, pre-agreed and tested. Either the responder holds standing read-only access to your cloud accounts and identity provider, or there is a documented path to grant it in minutes, tested at least once, that does not depend on a single person being awake. The common failure is a path that requires approval from the same administrator whose account is the one suspected of being compromised.

Out-of-band communications. If your email and chat run on the platform that is compromised, you cannot coordinate the response on them, and you should assume an attacker with mailbox access is reading the incident thread. Agree now on a fallback channel and make sure the phone numbers in it are personal rather than routed through the corporate system.

Log retention long enough to answer the question. Most breaches are discovered well after the initial access. If your cloud audit logs retain ninety days and the intrusion started five months ago, the investigation cannot establish what was taken, and cannot establish what was not taken either. That second one matters, because being unable to rule out data access often forces a broader notification than the facts would have required. Extending retention costs a modest amount monthly and is the highest-value thing most companies can do before an incident.

A current asset and data map. Not a perfect inventory. A list of what runs where, which datastores hold personal information, which vendors hold data on your behalf, and who owns each system. Responders spend their first hours building this if you have not, and they build it worse than you would have, because they are reconstructing it under pressure from people who are also under pressure.

What the first hour looks like when you have no retainer

Plenty of readers will find this article mid-incident. In that case, order matters more than speed.

Preserve before you fix. The instinct is to rebuild the compromised server or reimage the laptop, and doing so destroys the evidence needed to determine scope. Isolate the system at the network level rather than powering it off, take a snapshot of the disk and, if you have the capability, of memory. Export the relevant logs to somewhere outside the affected environment, because logs sitting inside a compromised platform can be deleted by whoever compromised it.

Contain identity before infrastructure. In most cloud-first breaches the foothold is a credential, so revoke active sessions and rotate credentials for the affected accounts, including any API keys and service accounts they could reach. Rotating a password without revoking the existing session tokens is the most common containment mistake we see, and the attacker simply stays logged in.

Then call your insurer, your counsel and a responder, in that order, and start a written timeline with clock times. That timeline becomes the backbone of every subsequent conversation with regulators, customers and the insurer, and nobody can reconstruct it accurately three days later from memory.

What a real engagement costs, and what drives it

We will not quote a number for someone else's incident because the range is genuinely enormous, but the drivers are predictable. Time to discovery is the biggest: an intrusion found in a day is a contained problem, and one found after months is an investigation across whatever log history you happen to have. The second is estate size and heterogeneity, because every additional platform is another set of logs to collect and interpret. The third is whether data was exfiltrated, since that shifts the work from technical response to legal analysis, notification and, quite often, an extended period of customer management. The fourth is your own readiness, which is the only one of the four you can change today.

Preparation is cheap by comparison. A tabletop, a runbook, extended log retention and a tested access path together cost a small fraction of one incident and reduce the duration of every one that follows.

Keeping the arrangement warm

A retainer decays quietly. The responder holds access to a cloud account you deprecated, the escalation contact left in March, the runbook describes a monitoring tool you replaced, and none of it surfaces until the night it matters. Twice a year, spend an hour confirming three things: that the access path still works when someone actually tries it, that every name and phone number in the escalation list belongs to a current employee, and that the environment described in the onboarding pack resembles the environment you now run. Ask your provider to initiate that check rather than waiting for you, and treat a provider who never asks as one you should be reviewing at renewal.

When a retainer is the wrong purchase

If you have no centralised logging, no managed endpoints and no single identity provider, a retainer buys you a responder who will arrive and tell you they cannot see anything. Spend the money on visibility first: consolidate identity, get endpoint detection deployed, and get audit logs into somewhere queryable with a retention window measured in months. Then buy the retainer. Doing it in the other order is common and it wastes both the fee and the first day of your incident.

If your carrier mandates a panel firm and your environment is small and well instrumented, you may already have adequate incident coverage through the policy. Confirm the response times in the policy documentation rather than assuming, but if they are contractual and the panel firm is competent, a second retainer may be duplicated spend better put into detection or into testing whether the attack paths exist in the first place.

If you have a managed service provider running your infrastructure, check what their contract already commits them to. Some MSP agreements include meaningful incident support and some include a best-efforts sentence that means nothing. Read it before buying alongside it, and if it is the second kind, say so to your MSP rather than quietly working around them.

And if you are a five-person company pre-revenue with no customer data of consequence, the honest answer is that a documented plan, a phone number for a firm you have spoken to once, and good backups you have actually restored from will serve you better than a retainer fee this year. Revisit it when you hold customer data you would have to notify people about. When you do reach that point, the retainer and the readiness work are the same conversation, and we would rather have it with you early than at two in the morning.

Before you need it. Incident response on retainer means the contracts, the access and the runbooks already exist when the pager goes off.

See how a retainer worksOr talk about a retainer

Before you go

Want the rest of this by email?

If this was useful, I send a few short notes on incident response. Unsubscribe in one click, and replies reach me directly.

From Jacob Masse, principal of traztech. No spam, unsubscribe in one click.

Want a second opinion on where you stand?

We run SOC 2, ISO 27001 and the rest of the compliance stack for startups and SMEs, and the security testing that sits behind it. The first call is free, and we will tell you if you are not ready to start yet.

Book a free call

Track record

Who is actually doing the work

5
Published CVEs, including a CVSS 9.1
76
Controls taken from nothing to a passed SOC 2 Type II
Zero
Exceptions on that Type II report
20+
Penetration testing engagements delivered

Published vulnerability research

Five published CVEs. CVE-2024-45163 (CVSS 9.1) is a flaw in the Mirai botnet itself, which gave defenders a way to shut down attacker infrastructure. CVE-2026-42626 takes HP ENVY 5000 printers offline from any unauthenticated device on the same network.

A SOC 2 Type II built from nothing

At Humera, a venture-backed US security company, Jacob built the compliance programme in-house from nothing: no report, no policies, no documented controls. It ended in a Type II attestation across 76 controls with zero exceptions, on a team of 15.