Direct answer: An incident response retainer buys you a relationship, an access path and a defined process before you need them. The response time in the agreement matters less than what happens inside it: who picks up, what they are authorised to do, whether they already have access to your environment, and whether the scope covers your actual failure modes. Ask what an engagement looks like in hour one, not what the acknowledgement target is.
What you are actually buying
During an incident, the expensive delay is almost never finding somebody technical. It is contracting, scoping and access. A firm you have never worked with needs an agreement signed, an NDA, a scope, a rate, and read access into systems you are currently unsure about. That process takes hours to days, and it happens while your team is trying to contain something.
A retainer front-loads all of it. The contract exists, the access paths exist, the runbook exists, and somebody on the other end already knows what your stack looks like. That is the product. The response time in the SLA is a promise about how fast the relationship activates, not about how fast the problem gets solved.
Questions worth asking
Who picks up, and what is their background? A rotation staffed by tier-one analysts reading a script is a different product from a named responder who has run breaches. Ask specifically who answers at 2am on a Saturday.
What are they authorised to do? Some retainers are advisory: they tell your team what to do. Some are hands-on within an agreed boundary. Both are legitimate and the difference matters enormously at the moment you need it. Get it in the scope in writing.
What access exists in advance? Read-only into cloud, identity provider and observability, established during onboarding, is the difference between working the problem in hour one and negotiating credentials in hour one.
What is out of scope? Deep forensics for litigation, malware reverse engineering, negotiating with a ransomware operator, and legal disclosure drafting are commonly excluded or subcontracted. Knowing which of those your retainer covers is not something to discover during an incident.
What happens between incidents? A retainer that only earns its money during a breach is poor value in the years you do not have one. Tabletops, runbook maintenance, escalation tree upkeep and post-incident improvements are what makes the quiet months worth paying for.
How do hours work? Included hours, rollover, and the rate once they are exhausted. An incident that runs a week will exhaust any included allowance, so the post-allowance rate is the number that actually matters for a bad month.
What the SLA does and does not promise
Acknowledgement time is when someone confirms they have your escalation. Time to hands-on is when work starts. These are different numbers and vendors are not always careful about which one they quote. Ask for both, and ask what the coverage actually is, because 24/7 for acknowledgement with business-hours engineering is a common shape.
Also worth checking: whether the SLA is contractual with a remedy, or a target. Most are targets. That is not scandalous, but it changes what the number means.
The compliance and insurance angle
Two audiences care about this beyond your own risk. Enterprise security questionnaires routinely ask whether you have a retained incident response capability, and answering yes with a documented SLA closes the row cleanly. Cyber insurers increasingly ask the same question at underwriting, and some require you to use a panel firm during a claim, which is worth checking against your retainer before you buy either.
SOC 2 and ISO 27001 do not require a retainer. They require an incident response plan that has been tested, which a tabletop satisfies and a retainer usually includes. If your primary motivation is the audit rather than the risk, a tested plan is the cheaper and more direct purchase.
When a tabletop is the better buy
If you have never run an exercise, start there. A one-day facilitated tabletop against a realistic scenario for your stack tells you whether your plan works, produces the test record your framework asks for, and gives you a much clearer view of what you would actually need on retainer. Most teams change their mind about the tier they need after one.
The pattern that works well is a tabletop first, then a retainer sized to what the exercise exposed, rather than a tier chosen from a table before anyone has tested anything.
What we do differently
We stopped selling incident response as a standalone retainer with SLA tiers, because the companies buying it were buying the rest as well: someone to keep the compliance programme true, answer questionnaires and own the security roadmap. Splitting that into separate agreements was tidy on an invoice and wrong in practice, since the same people do all of it and the context transfers.
It runs through one retainer now, scoped and priced per client, with the response commitment written into the agreement where the work needs one. A one-day tabletop is still a fixed-scope engagement with a published price, and it remains the right first step for most teams.
What hour one actually looks like
Ask any vendor to walk you through the first hour and listen for whether the answer is operational or promotional. A competent one describes something close to this.
You call or page the escalation number and reach a responder, not a ticket queue. Within minutes there is a bridge open and a short scoping conversation: what you saw, when, on what system, and what your team has already done. That last question matters more than it sounds, because well-meaning containment frequently destroys the evidence needed to answer the question your customers will ask, which is what was accessed. A responder will tell you to snapshot before you rebuild, to isolate a host at the network layer rather than powering it off where memory matters, and to stop rotating credentials until they have captured which ones were used and when.
Then comes triage on data you already have. Identity provider sign-in logs, cloud audit trails, EDR telemetry, mail flow rules. The first hour is rarely about deploying tooling. It is about establishing a timeline from what already exists, and the quality of that hour is determined almost entirely by decisions you made months earlier about what to log and how long to keep it.
Also worth asking: who runs communications. Somebody has to decide what goes to customers, when a regulator clock starts, and whether the legal team engages. A good retainer has an opinion about that split and does not assume your team knows.
What you have to bring for a retainer to work
A retainer accelerates a response, it does not manufacture the conditions for one. The things that determine whether hour one is productive are all on your side.
Log retention. This is the single biggest determinant of investigation cost and of whether you can make a defensible statement about scope. Intrusions are frequently discovered weeks after initial access. Thirty days of retention on cloud audit logs and identity events means an investigation into a 45-day-old compromise cannot establish what happened, and the honest report says so. A year of retention on authentication and cloud control-plane logs is the highest-value security spend most companies of this size are not making.
An asset and identity picture. Not a perfect CMDB, just a current answer to what cloud accounts exist, who has administrative access to each, which systems hold customer data, and what your external attack surface is. Responders will build this anyway. Building it during an incident costs hours at incident rates.
Out-of-band communications. If your incident bridge runs on the platform that has been compromised, you are briefing the attacker. Agree a fallback channel and a phone tree during onboarding and store it somewhere reachable when your identity provider is untrusted.
A contact tree with authority attached. Named people for the decision to take a service offline, the decision to notify customers, and the decision to spend money. Incidents stall on approvals more often than on analysis.
Onboarding is the part to inspect
The gap between a good retainer and a badge on your website is almost entirely in onboarding. Ask what happens in the first 30 days after signature, and expect specifics: an architecture and log-source review, read-only access provisioned and tested, an escalation runbook written against your stack, and a documented contact tree with authority levels.
Then test it. Call the escalation number on a Saturday and see who answers and how quickly. Ask for the access to be exercised once so you discover the broken IAM role now rather than during a compromise. Re-run that test after any material change to your environment, and at least annually, because credentials rot and rotations quietly break the access path you are paying to have in place.
Ask also how the relationship handles your growth. A retainer scoped against a two-account AWS footprint does not automatically cover the Azure tenant you acquired last quarter.
What a real incident costs beyond the retainer
The retainer fee is rarely the meaningful number. The costs that expand are analysis hours across a multi-week engagement, tooling deployed for the response such as an EDR agent rolled out to every endpoint under an emergency licence, forensic imaging and storage where litigation or regulatory exposure is likely, and the internal engineering time you will not be spending on product. A serious intrusion consumes weeks of your senior engineers, and that cost never appears on anyone's quote.
One structural point worth raising with your lawyer before you sign anything. In matters likely to end in litigation or regulatory action, the response is often engaged through external counsel so that findings sit under legal privilege. If that is the shape you want, the retainer needs to permit assignment or engagement through counsel, and your firm needs to be acceptable to your insurer's panel. Working this out at signature takes an email. Working it out mid-incident costs days.
When the money is better spent elsewhere
If your logging is thin, buy logging before you buy a retainer. A responder with 7 days of retention and no endpoint telemetry can help you contain, but cannot tell you what was taken, which is the question your customers, your regulator and your insurer will all ask. Extending retention and turning on the audit trails you already have available costs a fraction of a retainer and improves every future outcome.
If nobody has tested your plan, run the exercise first. The tabletop remains the better first purchase for most teams, and it usually changes what tier you think you need.
If you are a ten-person company with no regulated data, no enterprise contract requiring it, and an insurer who has not asked, a documented plan with your cloud provider's support tier and a named external contact you have actually spoken to may be sufficient for now. Buying a retainer to answer one questionnaire row is an expensive way to answer one questionnaire row.
Where a retainer earns its place: you hold data whose loss would be materially damaging, you have contractual notification obligations with tight clocks, you have a small team with no in-house responder, or you have already had an incident and discovered what the improvised version costs. If that is your position, the useful conversation starts with your log sources and your worst realistic day rather than with a tier table, and that is how we scope it under a retainer. If you are earlier than that, our security work and a short conversation will usually point you at the cheaper fix first.
Before you need it. Incident response on retainer means the contracts, the access and the runbooks already exist when the pager goes off.
See how a retainer worksOr talk about a retainer