Security

Real offensive depth

Testing and defence led by a published security researcher with five CVEs, including a CVSS 9.1 Mirai botnet kill-switch.

All security →
Compliance

Audit-ready, fixed scope

SOC 2, ISO, and the Canadian privacy stack, run end to end with an independent auditor.

All frameworks →
Resources

Learn the space

Original research, free tools, and plain-language guides on security and compliance, from a published security researcher.

Read the blog →
Security

The Real Cost of Downtime for a SaaS Startup

The standard way founders think about downtime is "lost revenue per hour." For a SaaS with $5M ARR, that is roughly $570 an hour. Annoying but absorbable. So why bother investing in reliability?

Because lost revenue is the smallest line on the downtime bill.

What downtime actually costs

Direct revenue loss. Real but usually small at startup scale. $500 to $5,000 per hour depending on ARR.

Customer trust. Churn after a major outage runs 2 to 5 percent above baseline for the following quarter, and is concentrated in your highest-value enterprise accounts. At $50K ACV, even a 1% churn bump is real money.

Support load. Every minute of customer-facing downtime generates 5 to 20 minutes of support work as tickets come in, you write status updates, you do follow-up communications, and you handle credit requests. A two-hour outage typically eats 20+ hours of support team time.

Engineering response. A serious incident pulls 3 to 8 engineers off their planned work for the duration of the incident plus several days of postmortem and remediation. That is a week of engineering capacity gone.

SLA credits. If you signed enterprise contracts with uptime guarantees (you almost certainly did), you owe credits proportional to time below SLA. These show up in next quarter's revenue line.

Sales pipeline damage. Prospects in active evaluation will see the status page or hear about it from your existing customers. We have seen 6-figure deals stall after a single major outage in mid-cycle.

Recruiting damage. Senior engineering candidates Google "your-company outage" before accepting an offer. A pattern of incidents is visible.

The math that actually matters

For a $10M ARR SaaS, a two-hour total outage costs roughly:

  • Direct revenue: $2,300
  • Support team: 20 hours × $75 = $1,500
  • Engineering time: 40 hours × $150 = $6,000
  • SLA credits: $5,000 to $25,000
  • Elevated churn (one quarter, 1% of base): $25,000
  • Pipeline impact: $50,000+ (one stalled deal)

Total realistic cost: $90,000 to $110,000 for a single 2-hour outage. The "lost revenue" line is less than 3% of that.

Before you need it. Incident response on retainer means the contracts, the access and the runbooks already exist when the pager goes off. See how a retainer works

What is actually worth investing in

Given the math, even modest reliability investments pay back fast. The high-leverage moves:

A real status page (Statuspage, BetterStack, or Instatus). Customers who can see status during an outage generate a third as many support tickets as customers who cannot.

Proper monitoring and alerting (Datadog, Grafana, or equivalent). Mean time to detect drives mean time to resolve directly. Detecting an issue in 2 minutes vs 30 minutes is the difference between a paper cut and a 6-figure incident.

An incident runbook and on-call rotation. When the first thing your team does during an outage is figure out who to wake up, you have already lost 20 minutes. A documented rotation and runbook cuts response time dramatically.

Postmortems with actual follow-up. Most teams do postmortems. Few of them assign action items with deadlines and follow up. The incidents you do not learn from will repeat.

What is not worth investing in (yet)

Multi-region failover before you have repeatedly hit single-region issues. Five nines of availability for a product where customers are fine with three nines. A dedicated SRE team before you have 20+ engineers. These are real practices, but they are usually the wrong investment for early-stage startups. Solve the basics first: backups, monitoring, and a deploy pipeline you trust, which is where our DevOps work starts.

Want a reliability audit?

We audit production infrastructure, identify the highest-leverage reliability investments, and help you implement them in priority order.

Book a review

Measure it for your own company, not from a table

The numbers above are a model, and models are useful for arguing budget. What actually changes decisions is your own version, built from four inputs you already have. Take your fully loaded engineering cost per hour from payroll rather than from salary. Take your support cost per ticket from your helpdesk data. Take your average contract value and your enterprise account count from the CRM. Take your SLA credit terms from the three largest signed contracts, not from the template you send out.

Then instrument the thing most teams never capture: how long incidents actually last, from first customer impact to full recovery, not from the alert firing to the fix deploying. Those are different numbers and the gap between them is usually where the cost lives. If your incident tracker records only the second one, your postmortems are systematically understating impact and your reliability budget is being argued from bad data.

Run this once a quarter over the incidents you actually had. Within two quarters you will have a defensible cost-per-incident-hour for your business, which is the number that wins the argument about whether to spend three engineering weeks on the deploy pipeline.

The outages that never show up as outages

Total outages are rare and memorable. Partial degradation is common and largely uncounted, and for most SaaS businesses it does more cumulative damage. Checkout works but takes eleven seconds. The API is up but the webhook queue is nine hours behind. Reporting loads for small accounts and times out for the three largest ones, which are the accounts that matter.

Uptime percentage hides all of this, because the health check passed the whole time. The customer experience was an outage; the status page said operational. That mismatch is expensive twice over: you take the churn and support cost anyway, and you also take the credibility hit from a status page that told customers nothing was wrong while they could not work.

Two fixes are worth more than any infrastructure investment here. Define availability from the customer's side, using synthetic transactions that exercise the workflows customers actually run, per major account tier. And set an explicit threshold at which degradation gets posted publicly, so the decision is made in advance rather than by whoever is on call and does not want to escalate.

Read your SLA before you price the risk

Founders assume the SLA exposure is the credit percentage. It usually is not. Credits in a standard enterprise agreement are capped at some share of the monthly fee, which for a $50,000 annual contract is a few thousand dollars in the worst month. Unpleasant, survivable.

The clauses that carry the real risk sit further down. Termination for chronic failure lets the customer exit without penalty if you breach the availability target in some number of consecutive or rolling months, which converts an operational problem into a lost renewal. Definitions of downtime decide whether degraded performance counts at all, and whether your own maintenance windows and your cloud provider's failures are excluded. Measurement authority decides whether your monitoring or the customer's determines whether a breach occurred, and customers increasingly insist it is theirs. Notification obligations impose a clock, often one or two hours, and missing it is its own breach independent of the outage.

Pull those four clauses out of your top ten contracts and put them on one page. Most teams discover they have committed to different targets for different customers with no way to tell operationally which customers are on which tier. Our SLA guide goes into how to write terms you can actually meet.

When the outage is a security incident

Availability incidents and security incidents share a cost model until the moment you suspect an attacker, at which point the economics change completely. Ransomware, a compromised credential in production, or a supply chain compromise all bring recovery costs plus notification obligations, forensic work, legal involvement, cyber insurance process and regulatory clocks under PIPEDA or a provincial equivalent.

The operational difference that catches teams out is that you cannot restore and move on. Restoring from backup into an environment you have not established is clean is how organizations get encrypted twice. Containment, evidence preservation and root cause come first, and they take days rather than hours. Budget a security-driven outage at an order of magnitude above an equivalent-length infrastructure outage, and read the first 24 hours of a ransomware incident before you need it.

The other difference is insurance. Cyber policies typically carry a business interruption waiting period, often eight to twelve hours, before loss accrues, and they require notification to the insurer's panel rather than to the responder you would have picked. If your plan says "call our usual firm", check the policy first, because using an off-panel responder can void the coverage you are counting on.

Downtime as an audit finding

If you carry the Availability criterion in a SOC 2, outages stop being purely commercial. The auditor is not testing whether you stayed up; they are testing whether the controls you described operated. That means incident tickets exist for the incidents that happened, escalation followed your documented path, customers were notified within your stated commitment, the postmortem was completed, and the remediation items were tracked to closure.

The exception that gets written is almost never "the system was down". It is that three of the eight sampled incidents have no postmortem, or that the runbook says notify within one hour and the record shows six. That is entirely a discipline problem and it is fixable for free. If Availability is in scope for you, treat the incident record as audit evidence from the first day of the observation window, not as something to reconstruct afterwards.

The outage you did not cause and still pay for

A large share of the incidents a small SaaS lives through originate somewhere else. Your cloud region, your payment processor, your identity provider, your email delivery vendor, your CDN. Fault is irrelevant to the bill. Your customers cannot work, your support queue fills at the same rate, your enterprise accounts ask the same questions, and the incident consumes the same engineering hours while your team watches someone else's status page and can do nothing.

What differs is your recovery of costs. Your vendor's SLA credit is calculated against what you pay them, which is a fraction of what you owe your own customers for the same window. If you resell an identity provider inside a $50,000 contract and pay a few hundred dollars a month for it, the credit you receive will not cover the credit you issue. That gap is unrecoverable and it is worth knowing its size before an incident rather than during one.

Three things reduce the exposure and none of them are expensive. Write down every third-party service that can take your product down, which is usually a shorter list than teams expect and always contains one dependency nobody had thought of. Decide in advance what degraded operation looks like for each, because a product that can run read-only through a payments outage keeps most of its customers working. And write the customer communication for a vendor-caused outage before you need it, because the instinct in the moment is to explain that it is not your fault, and that is the message enterprise customers remember worst. They bought availability from you, not from your vendor.

The vendor list has a second use. Auditors and enterprise security reviewers both ask for it, and a maintained dependency inventory with owners and criticality is most of what a vendor management control requires.

When outside help is not worth buying

Do not hire us for a reliability review if you already know what is broken. If your team can name the fragile subsystem and the reason it is fragile, an external assessment tells you what you already know at consultant rates. Spend the money on the engineering time to fix it. Where outside help earns its cost is when incidents keep coming from places nobody predicted, when a customer or auditor requires independent assurance, or when you want the incident response contracts, access and runbooks in place before the pager goes off rather than negotiated during a crisis. That last one is what a retainer exists for, and it is worth buying precisely because it is the only part of this you cannot arrange after the fact.

The line item nobody puts in the model: the people

Incident cost models count engineering hours at a blended rate and stop there. The expensive part is what a bad year of on-call does to retention. A rotation with too few people in it, alerts that fire without being actionable, and a pattern of overnight pages is a reliable way to lose senior infrastructure engineers, who are the hardest engineers to replace. Recruiting and ramping one replacement costs more than most of the reliability work you were arguing about funding.

There is a second effect that shows up in velocity rather than payroll. After a serious incident, teams get cautious. Deploys slow down, changes get batched into larger and riskier releases, and the batching causes the next incident. If your deploy frequency drops for a month after every outage, that is a real cost and you can measure it from your own pipeline data.

The cheapest counter is rotation hygiene rather than infrastructure. Enough people that nobody carries the pager more than one week in four, an explicit rule that a night page is followed by a late start, and a standing review that deletes or fixes any alert that fired without requiring action. Alert quality is the reliability investment with the shortest payback, because every noisy alert costs sleep and slows the response to the real one. Our guide to running an incident process covers how to structure the rotation and the severity levels around it.

Your outage history becomes a sales document

Once you sell to enterprises, uptime stops being an operational metric and becomes something you are asked to evidence. Procurement asks for twelve months of availability figures. Security reviewers ask how many severity one incidents you had and what changed afterwards. Prospects read the status page history during evaluation, and an empty history on a two year old product reads as a page nobody updates or a company that does not classify incidents.

The counter-intuitive part is that a visible incident with a well written postmortem does less damage than a suspiciously clean record. Buyers are not expecting perfection from a company of your size. They are checking whether you notice, communicate and learn. A postmortem that names what broke, what the customer experienced, and what shipped as a result is a credibility asset, and it costs an afternoon to write properly.

Where the recovery clock actually goes

Teams argue about reliability budget as though recovery time were a technical property of the system. Watch three of your own incidents closely and the shape is usually different. The fix, once the right person understands the problem, tends to be quick. The expensive hours sit either side of it.

The recurring delays are organisational. Detection lands in a channel nobody is watching at that hour. The engineer who was paged cannot reach production because the credentials sit in a vault held by someone asleep, or the break-glass path has never been used and nobody trusts it at 3am. The dashboards answer the question you had last quarter rather than the one in front of you, so twenty minutes disappear into working out which component is actually failing. Then a decision that carries risk, failing over, rolling back a migration, cutting a customer off from a degraded feature, waits for someone who feels authorised to make it.

None of that is solved by buying infrastructure, and all of it is cheap to solve. Put timestamps on five moments in your next few incidents: first customer impact, first human aware, first person with working access, cause identified, service restored. The longest gap is where the money should go, and for teams under fifty engineers it is almost always one of the first three. Circulate that finding internally, because it also settles the recurring argument about whether the next reliability quarter buys tooling or fixes process.

When the honest answer is to accept the downtime

Reliability spend competes with product, and for some businesses the right availability target is lower than founders like to admit. If customers use the product during working hours in one time zone, an overnight maintenance window costs nothing and buys real operational freedom. If your average contract is small and self-serve, a two hour outage generates support load and very little churn. Building for continuous availability in either case reduces a cost you do not have.

Be equally honest about what you promise. Committing to a target you cannot measure or meet converts an engineering problem into a contractual one, and the credits are the smallest part of that. Set the target you can evidence, state your maintenance windows plainly, and raise the commitment when the accounts justify it. If reliability failures are turning into customer security questions, or a buyer is asking for evidence you do not currently produce, that is the point where an outside view earns its fee rather than duplicating what your team already knows, and it is worth describing the situation before committing to a programme.

Before you need it. Incident response on retainer means the contracts, the access and the runbooks already exist when the pager goes off.

See how a retainer worksOr talk about a retainer

What we charge for this. The figures above are market ranges. Our own fixed-scope prices are on the pricing page, alongside every cost breakdown we have written.

Before you go

Want the rest of this by email?

If this was useful, I send a few short notes on security posture. Unsubscribe in one click, and replies reach me directly.

From Jacob Masse, principal of traztech. No spam, unsubscribe in one click.

Want a second opinion on where you stand?

We run SOC 2, ISO 27001 and the rest of the compliance stack for startups and SMEs, and the security testing that sits behind it. The first call is free, and we will tell you if you are not ready to start yet.

Book a free call

Track record

Who is actually doing the work

5
Published CVEs, including a CVSS 9.1
76
Controls taken from nothing to a passed SOC 2 Type II
Zero
Exceptions on that Type II report
20+
Penetration testing engagements delivered

Published vulnerability research

Five published CVEs. CVE-2024-45163 (CVSS 9.1) is a flaw in the Mirai botnet itself, which gave defenders a way to shut down attacker infrastructure. CVE-2026-42626 takes HP ENVY 5000 printers offline from any unauthenticated device on the same network.

A SOC 2 Type II built from nothing

At Humera, a venture-backed US security company, Jacob built the compliance programme in-house from nothing: no report, no policies, no documented controls. It ended in a Type II attestation across 76 controls with zero exceptions, on a team of 15.