Security

Real offensive depth

Testing and defence led by a published security researcher with five CVEs, including a CVSS 9.1 Mirai botnet kill-switch.

All security →
Compliance

Audit-ready, fixed scope

SOC 2, ISO, and the Canadian privacy stack, run end to end with an independent auditor.

All frameworks →
Resources

Learn the space

Original research, free tools, and plain-language guides on security and compliance, from a published security researcher.

Read the blog →
home / guides / audit evidence

What evidence auditors actually accept

Most audit pain is not caused by missing controls. It is caused by artifacts that do not prove what they are offered to prove. Here is what gets accepted, what gets bounced, and why.

Last reviewed September 2026 · by traztech, security & compliance for startups
Short answer

An auditor accepts an artifact when it shows what happened, when it happened, who did it, and that the set is complete. That is why a policy is not evidence that a control operated: it states an intention, not an occurrence. The strongest evidence is system generated and dated at the time, next is an export you can re-derive, and last is a screenshot, which is acceptable only when the system produces nothing else and only if it carries the date, the system chrome, the scope, and who took it. The most expensive category of failure is not weak evidence, it is evidence that had to exist at the time and does not, because that cannot be produced afterwards at any price.

3 tiers
point in time, needs assembling, period evidence
4 questions
what, when, who, and is it complete
No backfill
contemporaneous records cannot be made later

The four questions every artifact has to answer

An auditor is not looking at your evidence to admire the control. They are testing an assertion, and an artifact is accepted when it answers four questions without anybody having to explain it: what happened, when it happened, who did it, and whether the set in front of them is complete.

Almost every rejection traces back to one of those four. A screenshot with no date fails the second. A spreadsheet somebody assembled fails the fourth, because there is no way to tell it apart from an edited one. A configuration page showing the setting is correct today fails the second when the question was about a period that ended in March. A ticket with no assignee fails the third.

Getting this right is worth real money, and not in an abstract way. Fieldwork is billed against a plan, and every artifact that gets bounced generates a follow-up request, a round trip, and a delay. A program where the auditor accepts the first submission on most items finishes on schedule. A program where a third of items come back finishes late and consumes weeks of internal time, which is the part of the cost nobody budgets for. Our SOC 2 cost guide covers where the money actually goes, and internal time is consistently the most underestimated line.

On the engagements we run there is a step between collecting evidence and handing it over, where every artifact is reviewed against what the audit firm will accept before the audit firm sees it. That step exists because it is far cheaper to reject your own evidence than to have it rejected during fieldwork.

Why a policy is not evidence that a control operated

This is the single most common misunderstanding, and it costs teams more time than any other. A policy is evidence of design. It says what is supposed to happen. It says nothing about whether it happened, and a Type II opinion or an ISO certification is largely about whether it happened.

The distinction the frameworks draw is between design and operating effectiveness. Design asks whether a control, if it operated as described, would achieve the objective. Operating effectiveness asks whether it actually operated, consistently, across a period. A Type I report and the design side of an ISO audit deal mostly with the first. A Type II report and the surveillance cycle deal with the second, and the second is where the evidence burden lives.

So for each control you generally need three things and most teams have one. You need the statement, which is the policy or procedure. You need the configuration, which shows the mechanism exists and is set correctly. And you need the occurrence record, which shows it happened on specific dates to specific things. A password policy document, a screenshot of the identity provider enforcing complexity and MFA, and an export showing MFA enrolment state for every user at period end is the full set. Only the first of those is a policy, and a submission consisting of the policy alone is the most frequent reason an item gets bounced.

The same logic explains why an approved and unread policy is weak even as design evidence. If the control depends on people following the policy, the auditor will want acknowledgements with dates, and if the policy changed mid-period, acknowledgements against the version that was current at the time. A policy set nobody has acknowledged since 2024 does not evidence a workforce that knows the rules.

The three tiers of evidence, and why the distinction matters

It is useful to sort every requested item into one of three tiers before you start, because the tiers tell you what can be done today, what needs a week, and what cannot be done at all if you have left it too late. We tier every request list this way on delivery, and it is the fastest way to see how much trouble a program is in.

Tier 1 is point in time. You either have it or you do not, and if you have it you can send it today. Policies, network diagrams, the asset register, the vendor list, contracts, standards, procedures, configuration exports. There is no assembly and no waiting. A program with Tier 1 gaps has a documentation problem, which is annoying and solvable in weeks.

Tier 2 needs assembling. It does not depend on a period of operation, but somebody has to build it: a reconciliation, a data flow map, a population export enriched with fields from a second system, a summary of something across accounts. Days rather than hours, and it is where most of the drudgery sits.

Tier 3 is period evidence. It only exists if the control has actually been running for a while. Twelve months of access reviews, a quarter of change tickets, a year of vulnerability scans, monthly reports to management, a visitor log with entries in it, training completions across the period. This is the tier that determines whether you have a report this year or next year, and no amount of effort in the final month produces it.

Tiering the list is the first useful diagnostic in any readiness assessment, because the Tier 3 column tells you the truth. A company with thirty Tier 1 gaps and no Tier 3 gaps is four weeks from an audit. A company with three Tier 1 gaps and twenty Tier 3 gaps is at least a full observation window away, however good its documentation looks.

Cadence, and the records that cannot be created later

Every recurring artifact carries two properties, and the second one is the one that ends audits. The first is cadence: how often it has to be produced across the period. The second is whether it is contemporaneous: whether it can be created after the fact at all.

Some records only exist if they were made at the time. A visitor log for a month nobody kept one cannot be produced afterwards. An access review dated last quarter cannot be run today. A restore test that was not performed in March cannot be performed retroactively in September. A monthly management report series assembled the week before fieldwork reads exactly like what it is, because the file timestamps, the formatting consistency and the content all give it away, and an auditor who spots it will reasonably start doubting everything else you handed over.

The practical consequence is severe and worth stating plainly. If a contemporaneous artifact is missing for a period inside your observation window, that period is permanently unevidenced. There is no catching up. The remaining options are a shorter window, a later report date, or a qualified opinion with the exception described in the report. Which of those is least bad depends on what you promised customers, and it is a conversation to have with the audit firm early rather than at the readout.

This is why the sequencing of a readiness engagement matters more than the effort put into it. Work that unblocks evidence collection comes before work that is merely important, because a control that starts producing dated artifacts in month one has eleven months of evidence by the time the window closes, and the same control started in month nine has three. We deliberately order a remediation plan by what blocks evidence collection rather than by severity for exactly this reason.

The list of records that are effectively always contemporaneous is worth memorising: access reviews, change tickets, incident records, tabletop and drill records, restore tests, training completions, policy acknowledgements, board or management review minutes, internal audit records, visitor logs, disposal logs, joiner and leaver records, and any log, alert or scan output. If it appears in that list and you are inside a window, produce it on schedule or lose the period.

System generated, exported, or screenshotted

There is a strength ordering to evidence forms, and knowing it tells you what to reach for first.

System generated records are strongest. An audit log entry, a ticket with its own history, a scan report produced by the scanner, a pipeline record, a signed agreement with a certificate trail. These carry their own dating and their own authorship, they are hard to fabricate without leaving traces, and the auditor does not have to trust your account of them. Where a control can be evidenced this way, it should be.

Exports are next. A CSV from a console, a report from the identity provider, an output from a CLI command. These are strong when they are re-derivable, which means recording where the export came from, what filter was applied, when it was taken, and how many rows it contained. An export whose provenance is written down is nearly as good as a system record. An export pasted into a spreadsheet, reformatted and columns removed, is a document you made.

Screenshots are weakest and sometimes unavoidable. Plenty of SaaS admin consoles will not export the thing the auditor wants, and a screenshot is the only available artifact. That is fine, and auditors accept them constantly. What they do not accept is a cropped rectangle showing a toggle in the on position with no context. A usable screenshot shows the full browser or application window including the URL or system identifier, the account or tenant it was taken in, the system clock or an in-page date, the full scope of what is being shown rather than a filtered subset that hides the exceptions, and it is accompanied by a note saying who took it, on what date, and what it is offered to prove.

Assertions are not evidence. A tick in a tracker, an email saying the thing was done, or a spreadsheet cell marked complete is the compliance owner vouching for someone else's work, which is the thing under test. These are useful for managing the program and worthless as evidence, and confusing the two is why some registers look green all the way to fieldwork.

Dating and completeness

Dating sounds trivial and it is the most common single defect. An artifact needs a date that ties it to the period under audit, visible in the artifact itself rather than in the filename or the folder, because filenames are typed and system dates are not.

Watch for the specific case of the evidence that is true now and says nothing about then. A configuration screenshot taken in September proves the setting is correct in September. If the audit period ran from January to June, that artifact is out of scope for the period, and this catches people constantly because the control genuinely was operating the whole time. Where the system keeps configuration history or change logs, use those. Where it does not, the answer is to capture configuration state periodically during the window rather than once at the end, which is a five-minute recurring task that saves an argument later.

Completeness is the other half. An auditor needs to know that what they are looking at is the whole set rather than a favourable slice. That means row counts on exports, an explicit statement of the filter applied, and no silent exclusions. If you export change tickets for the period and filter out the ones that were rejected, you have produced a misleading population, and if the auditor finds the rejected ones by another route the conversation changes character entirely.

Where an export genuinely has to be filtered, state the filter in writing next to the artifact. "All production change tickets closed between 2026-01-01 and 2026-06-30, excluding tickets in the Documentation project, 412 rows" is a complete and honest population definition. The exclusion is defensible because it is stated. An unstated exclusion is not defensible even when it is reasonable.

Population and sample

Most operating effectiveness testing is sampling, and sampling has a structure worth understanding because it determines what you need to hand over.

The auditor first establishes the population: every occurrence of the thing in the period. Every change, every joiner, every access review, every incident, every deployment. Then they establish how many times the control should have operated, then they select occurrences and test each one. The population comes from you, and this is the leverage point. If your population is wrong, the sample is drawn from a wrong population, and that is a completeness exception which is more serious than a control exception because it undermines every conclusion drawn from that population.

Sample sizes vary by firm, by how often the control operates, and by assessed risk. As a rough shape: a quarterly control across a year has four occurrences and commonly two are tested; a monthly control has twelve and commonly two to four are tested; a control that operates many times, like change management, gets a sample often somewhere between twenty and sixty depending on population size and risk. A control with an exception in the prior year attracts a larger sample. Treat these as orientation rather than as a rule, and ask your audit firm for their approach during planning rather than guessing.

Where you sample internally, for example when running your own testing before an audit, the property that makes it defensible is that it can be re-derived. Record the population size, the selection method, and the seed or the rule, so the same population and the same seed always produce the same selection. That is the difference between a sample an auditor accepts and one they replace. "We picked a few" is not a method, and it is a phrase we hear more often than it should be said.

One more thing about populations: get them early. The core populations, meaning the roster, the joiners and leavers, the vendor list and the change tickets, are the lists everything else samples from, which is why they are requested in the first days of an engagement rather than partway through. A program that cannot produce a clean roster on request is going to struggle with every population downstream of it.

The four things an auditor actually does

Understanding the procedures explains why some evidence is treated as stronger than others, and why the same fact sometimes has to be shown two ways.

Inquiry is asking someone. It is the weakest procedure and never sufficient on its own, which is why "we told them how it works" is not a completed test. It is, however, where the walkthrough happens, and a walkthrough that contradicts the documentation is how a lot of problems get discovered. Expect the auditor to ask the engineer, not the compliance owner, and expect the answers to be compared.

Observation is watching something happen. Useful for physical controls and for processes, and it is inherently point in time: watching a badge reader work today says nothing about March, which is why observation is usually combined with something else.

Inspection is examining a record, and it is the bulk of the work. Most of what you hand over is inspected. This is where dating, completeness and provenance decide whether the item passes.

Reperformance is the auditor doing it themselves. They pull the current user list and compare it to your reviewed population. They recalculate a metric. They attempt an action they expect to be blocked. Reperformance is the procedure most likely to produce a surprise, because it does not rely on anything you prepared. The defence is to reperform your own controls before handover, which is the same principle as reviewing your own evidence before submitting it.

The traits that get evidence rejected

From evidence review across readiness engagements, these are the defects that come up repeatedly. Each one is cheap to fix at collection time and expensive to fix during fieldwork.

Naming, filing, and the evidence register

A collection of correct artifacts in a badly organized folder still costs you fieldwork time, because the auditor cannot find things and asks, and every ask is a round trip. The register is the difference between a body of evidence and a shared drive.

Name files so that the artifact identifies itself without being opened: what it is, what it covers, and the date or period. "access-review-aws-prod-2026-Q2.csv" is self-describing. "export (3) final v2.csv" costs somebody a minute every time it is touched, and there will be four hundred files.

The register itself needs one row per required artifact carrying the control or controls it satisfies across every framework you are running, the owner, the cadence, whether it is contemporaneous, the period it covers, the current state, and where the file is. Collect once and map to many is the entire economics of running two frameworks at the same time: one dated access review record satisfies the logical access expectations of both SOC 2 and ISO 27001 without being produced twice.

Grooming matters before handover. Duplicates merged, near-identical items collapsed onto one artifact, names made consistent, and a state set on every row so nothing is ambiguous. The version of the register that is useful for running the program and the version that is ready to hand to an auditor are not the same document, and the gap between them is a real piece of work that is worth budgeting for rather than discovering. Keeping the register alive between audits is covered in keeping evidence fresh.

Reviewing your own evidence before the auditor sees it

The highest-leverage hour in a readiness program is spent rejecting your own evidence. Take each artifact, ask the four questions, and be unkind about the answers. What happened, when, who, and is this the whole set.

A workable review pass takes each item and asks: does the date on the face of it fall inside the period; is the scope on the face of it the whole population or a stated subset; is there a name attached to any decision it contains; is it in the strongest form the system can produce, or did somebody screenshot something that could have been exported; and does it contradict anything else in the package. That last check is the one people skip and it is where the real risk sits, because contradictions between artifacts are what turn a single finding into a scope conversation.

Run the reperformance checks yourself too, since they are the ones most likely to surprise you. Pull the current user list and compare it to the last reviewed population. Compare the leaver list to every system export. Compare the change population to the deployment history. Compare the vendor register to the expense records. Each of those takes minutes and each of them regularly finds something.

Handing over a package that has been through this pass changes the character of fieldwork. Requests get turned around inside a couple of days rather than a couple of weeks, the auditor stops sending follow-ups on the same items, and the engagement finishes on the planned date. That is worth more than it sounds, because a delayed report is usually a delayed deal.

When the evidence does not exist

Sometimes the honest answer is that the control did not operate, or operated without leaving a record. What you do next matters more than the gap itself, and the instinct to paper over it is the wrong one every time.

If the control operated but left no artifact, and the underlying record can be reconstructed from a system that was recording anyway, reconstruct it and say plainly that you did. An access review reconstructed from the ticket history and the identity provider logs, presented as a reconstruction, is a weaker artifact honestly labelled. The same reconstruction presented as a contemporaneous record is a misrepresentation, and if it is discovered the auditor will reasonably reconsider everything else.

If the control genuinely did not operate, the professional move is to raise it yourself, open a corrective action with an owner and a date, fix the mechanism so the next period produces the record automatically, and show the improvement. A self-identified exception with a corrective action and evidence of remediation reads very differently from one the auditor found. Under ISO in particular, a management system that finds and fixes its own nonconformities is what the standard is asking for, so a corrective action trail is closer to the point than a clean sheet.

And if a contemporaneous record is missing for a period inside the window, deal with it as a scoping decision rather than an evidence problem. Shorten the window, move the date, or accept the exception. Our post on why SOC 2 audits fail covers what those choices look like in practice, and what an exception in a report actually means is worth reading before assuming a qualified opinion is fatal. It usually is not.

Building a program that produces evidence as a by-product

The last point is the one that decides whether year two is easier than year one. Evidence that is collected is a chore that competes with everything else and eventually loses. Evidence that is produced automatically by work people were doing anyway survives.

The pattern is to move controls into systems that keep their own records. Approvals in the ticketing system rather than in email. Access changes through a request workflow rather than a direct message. Deployments through a pipeline that records who and when. Reviews in a tool that timestamps the decision. Each of those changes moves an artifact from something you have to remember to make into something that exists because the work happened.

What cannot be automated is the judgement: the review decisions, the risk acceptances, the exception handling, the management review. Those need a person, a date, and a record, and they are the items to protect in the calendar because they are the contemporaneous ones. Everything else should trend toward being generated.

The measure of whether you are winning is simple. Ask what the next audit would cost in internal hours if it started tomorrow. If the answer is that most artifacts already exist and the work is assembly, the program is healthy. If the answer is that everything has to be gathered, the program is a project that runs annually, and the second year will cost roughly what the first did. Our post on maintaining SOC 2 in year two goes into what changes and what does not.

Evidence forms, ranked by how well they hold up

Form of evidence Strength What it has to carry, and where it fails
System audit log entry Strongest Timestamp, actor, and action, produced by the system itself. Fails only if the export is filtered without saying so or the retention window is shorter than the audit period.
Ticket with its own history Strong Requester, approver, dates, and state transitions held by the tool. Fails when approvals happened in chat and were pasted in afterwards, or when the assignee field is empty.
Scanner or pipeline report Strong Generated by the tool with its own run date and scope. Fails when only the summary page is provided, or when the scope of the scan does not match the scope of the audit.
Console export with provenance Strong The raw file plus where it came from, the filter applied, the timestamp and the row count. Fails when it has been opened, reformatted and saved, because that is now a document you made.
Signed agreement Strong Both signatures, both dates, and the executed version rather than the draft. Fails when marked executed with no countersigned copy on file, which is the most common vendor register defect.
Generated record of a performed check Good A dated, itemised record naming who performed it, with a result against each item. Fails when it is undated, unattributed, or when failed items were not carried into a remediation register.
Meeting minutes or a dated memo Good Date, attendees, what was discussed, decisions with owners. Fails when the whole series was written in one sitting, which the formatting and file timestamps make obvious.
Screenshot Acceptable, weakest Full window, URL or system identifier, tenant or account, visible date, unfiltered scope, plus a note of who took it and what it proves. Fails cropped, undated, or filtered to hide exceptions.
Third-party report on a vendor Depends The full report, in date, with the exceptions section read and the complementary user entity controls acted on. Fails when only the cover page or the certificate is on file.
Policy or procedure document Design only Approved, versioned, dated, with acknowledgements. Evidences intent, never occurrence. Submitting it where operating evidence was requested is the single most common bounce.
Tracker tick or email assertion Not evidence Someone stating that something was done. Useful for running the program, worthless for testing it, because the assertion is the thing under test.
Reconstructed record Weak, if labelled Rebuilt after the fact from systems that were recording anyway. Acceptable only when presented honestly as a reconstruction. Presented as contemporaneous it is a misrepresentation.

Frequently asked

Are screenshots acceptable as audit evidence?

Yes, and auditors accept them every day, because plenty of systems will not export the thing being asked about. What gets rejected is a screenshot that has been cropped to a toggle with no context. A usable one shows the full window including the URL or system identifier, the account or tenant, a visible date from the system rather than from the filename, and the whole scope of what is being shown rather than a filtered view that hides the exceptions. Add a short note saying who took it, when, and what it is offered to prove. Where the system can export instead, export instead, because an export is stronger and takes the same amount of time.

Why was our policy rejected as evidence?

Because a policy is evidence of design, not of operation. It states what is supposed to happen. A Type II opinion and an ISO certification are largely about whether it happened, which needs an occurrence record: the dated review, the ticket, the log, the export. For most controls you need three things, and teams usually have one. The statement, which is the policy. The configuration, which shows the mechanism exists. And the occurrence, which shows it ran on specific dates. Submitting the policy where the occurrence was requested is the most common single reason an item comes back.

Can we recreate evidence we forgot to collect?

It depends entirely on whether the record was contemporaneous. If a system was recording anyway and you can reconstruct from its logs, you can produce a reconstruction and label it as one, which is weaker evidence honestly presented. If the artifact only exists because a person made it at the time, such as an access review, a restore test, a tabletop record or a visitor log, then no. That period is permanently unevidenced and the options are a shorter observation window, a later report date, or an exception in the report. What you must not do is produce it now and present it as though it was made then, because the file timestamps and formatting usually give it away and the consequence is that the auditor stops trusting the rest of the package.

How much evidence does an auditor actually sample?

It depends on how often the control operates, the assessed risk, and the firm. A quarterly control over a twelve-month period has four occurrences and commonly two are tested. A monthly control has twelve and commonly two to four are tested. A control that operates many times, such as change management, typically gets a sample somewhere in the range of twenty to sixty depending on population size. A control with an exception in the prior year attracts more. Ask your audit firm about their sampling approach during planning, because it tells you exactly how much evidence you need to have in retrievable shape.

Does compliance tooling collect evidence that auditors accept?

For the parts it can reach, generally yes. Automated collection from connected systems produces dated, consistently formatted artifacts and it removes most of the manual screenshotting, which is a real improvement. The limits are worth knowing. The tool only sees systems it is connected to, so anything outside the integrations is still manual. It cannot judge whether the population it pulled is the complete one. And it cannot produce the human artifacts, which are the review decisions, the risk acceptances, the tabletop records and the management reviews, and those are precisely the contemporaneous ones that cannot be recovered if missed.

What is the difference between design and operating effectiveness?

Design asks whether the control, if it operated as described, would achieve its objective. Operating effectiveness asks whether it actually operated, consistently, across a period. A SOC 2 Type I and the documentation stage of an ISO audit deal mostly with design, which is why they can be completed relatively quickly. A Type II and the ongoing ISO cycle deal with operation, which requires evidence spread across the whole period and cannot be compressed. That difference is the entire reason a Type II costs more and takes longer than a Type I.

What happens if we hand over evidence that contradicts other evidence?

It escalates. A single weak artifact is a follow-up request. Two artifacts that cannot both be true is a credibility question, and the auditor responds by widening the sample or testing adjacent controls. Typical contradictions are a leaver who still appears in a current user export, a production change with no corresponding ticket, an incident mentioned in one document and absent from the incident register, and a vendor on the expense records and not in the vendor register. Run those comparisons yourself before handover, because each takes minutes and each regularly finds something.

How should evidence be named and organized?

So that each artifact identifies itself without being opened: what it is, what it covers, and the date or period. Behind that, keep one register with a row per required artifact, carrying the controls it satisfies in every framework you run, an owner, the cadence, whether it is contemporaneous, the period covered, its current state, and the file location. The point of mapping one artifact to several frameworks is that a dated access review record satisfies the logical access expectations of both SOC 2 and ISO 27001 without being produced twice, and that reuse is most of what makes running two frameworks cheaper than running two programs.

Related

Walk every control yourself

traztech Workspace has every control of whichever frameworks apply to you, written in plain English, with somewhere to attach the proof. Free to use, with no card and no trial clock.

No credit card, no trial clock, no locked features. We make money when someone wants help closing the gaps, not from the Workspace.

traztech Workspace Other GRC platforms
Licence cost $0. Free forever, no card, no paid tier $7,500 to $50,000 a year, on an annual contract
Control library, evidence register, policy templates, risk register, vendor questionnaires, readiness scoring Included Included
What it costs inside an engagement with us $0. You need a workspace either way Unchanged. The subscription sits on top of the fee
What it does to your audit quote $11,000 off a five-figure quote on one engagement, for a documented readiness position Nothing. The audit firm prices your readiness, not your tooling

Pricing in the right column is what compliance automation platforms are publicly reported to charge; none of them publish a number, so treat it as a range rather than a quote. The $11,000 came off the audit firm's own number once the readiness position was documented (the engagement). Where a paid platform is the better buy, and the fuller comparison, is on the Workspace page.

Want your evidence reviewed before the auditor sees it?

We check every artifact against what the audit firm will accept, so items get accepted on first submission instead of generating a round trip in the middle of fieldwork.

Book a strategy call

Want the human version?

Get Jacob's take, by email

Jacob sends a few short, practical notes on getting security and compliance right without the months of pain. No fluff, unsubscribe in one click. Reply anytime; it reaches him directly.

From Jacob Masse, founder of traztech. No spam, unsubscribe in one click.

Track record

Who is actually doing the work

5
Published CVEs, including a CVSS 9.1
76
Controls taken from nothing to a passed SOC 2 Type II
Zero
Exceptions on that Type II report

Published vulnerability research

Five published CVEs. CVE-2024-45163 (CVSS 9.1) is a flaw in the Mirai botnet itself, which gave defenders a way to shut down attacker infrastructure. CVE-2026-42626 takes HP ENVY 5000 printers offline from any unauthenticated device on the same network.

A SOC 2 Type II built from nothing

At Humera, a venture-backed US security company, Jacob built the compliance programme in-house from nothing: no report, no policies, no documented controls. It ended in a Type II attestation across 76 controls with zero exceptions, on a team of 15.