Skip to content
Compliance

The audit evidence checklist for a penetration test

Most people buy a penetration test because somebody asked them for one - an auditor, a customer, a QSA. This is the checklist for making sure what you buy is what they will accept: what each framework actually requires, what belongs in the evidence pack, and the three places this goes wrong late enough to move an audit date.

11 min read

Key takeaways

  • Only PCI DSS writes the requirement down. SOC 2 and ISO 27001 do not mandate a penetration test at all - auditors ask for one because it is the most convenient evidence for criteria written in terms of outcomes.
  • The single most common gap is not the test. It is the absence of evidence that findings were fixed and verified, which PCI DSS requires explicitly at 11.4.4 and every other framework expects in practice.
  • A report with no dated testing window, no named scope and no methodology is hard for an assessor to accept whatever it found, because they cannot establish what it covered.
  • Scope has to match the boundary the framework cares about - the cardholder data environment, the system described in your SOC 2 description, the ISMS scope. A test of the wrong boundary is a real test and useless evidence.
  • Book so the retest lands before your evidence deadline, not the test. Remediation time is the part people forget to budget.

What each framework actually requires

Start here, because the four are not equivalent and treating them as one requirement is how people over-buy or under-deliver. One of them names penetration testing and sets a cadence. The other three do not mention it as a mandatory control at all, and ask for outcomes that a penetration test happens to be a convenient way of evidencing.

What the standard says, and what an assessor asks for in practice
FrameworkWhat the text requiresWhat is usually asked for
PCI DSS v4.0Requirement 11.4 requires a defined methodology (11.4.1), internal and external penetration testing at least once every 12 months and after significant change (11.4.2, 11.4.3), and that exploitable vulnerabilities found are corrected and testing repeated to verify the correction (11.4.4).The report, the methodology it followed, and evidence of the repeat testing at 11.4.4. This is the one framework where the retest is written into the requirement rather than being good practice.
SOC 2The Trust Services Criteria do not require a penetration test. They require that vulnerabilities are identified, evaluated and remediated, and that controls are monitored - CC4 and CC7 in the common criteria.A recent test report is the usual evidence offered for those criteria, together with what you did about the findings. An auditor may accept other evidence; most ask for this because it is the cleanest.
ISO/IEC 27001:2022Annex A.8.8 requires that technical vulnerabilities are identified, evaluated and addressed. Penetration testing is not named as a mandatory control anywhere in the standard.A test report as evidence that A.8.8 is operating, plus the record of what was done with the findings. Certification bodies vary in how much they want to see.
HIPAA Security RuleSection 164.308(a)(8) requires periodic technical and non-technical evaluation. Section 164.308(a)(1)(ii)(A) requires a risk analysis. Neither names penetration testing.Technical testing evidence supporting the evaluation, and a demonstrable link from what was found to your risk analysis and remediation.

The sentence to keep in mind

A penetration test is evidence that may support an assessment. It is not a certification or an attestation, and any provider who tells you their report certifies you against a standard is either being loose with language or does not understand the standard. Your assessor decides what satisfies it.

Before you book: what to settle

Every item here is cheap to answer now and expensive to answer late. The last two are the ones that move dates.

  1. Who is asking, and for what. An auditor mid-examination, a customer security questionnaire and a QSA all want different things from the same engagement. Ask for the request in writing and read what it actually says.
  2. Which boundary has to be covered. The cardholder data environment, the system described in your SOC 2 system description, the scope of your ISMS. A test of the wrong boundary is a real test and useless as evidence.
  3. What counts as a significant change since the last one. PCI names this explicitly; if you have re-platformed, added an authentication provider or moved a component, the previous test may not carry.
  4. Whether internal as well as external testing is required. PCI requires both. The others do not distinguish, but an auditor looking at a system with an internal trust boundary often asks.
  5. The date the evidence is needed, not the date the test can start. Then work backwards through remediation and retest, which is where the real time goes.
  6. Who will own the findings internally. A report delivered to somebody with no authority to schedule engineering work is how a finding is still open at the next audit.

Budget for remediation, not just the test

A five to ten day engagement is the short part. Fixing what it finds and having the fix verified is what determines whether you have evidence by your date. If the deadline is six weeks away, book now rather than in a month.

What the report has to contain

Assessors differ, but the items below are what makes a report usable as evidence rather than merely interesting. The first group exists so a reader can establish what was examined; without them a report cannot support a claim about coverage however good the testing was.

Establishing what was covered

  • A dated testing window, not just a publication date. An assessor needs to place the work inside the period under examination.
  • The scope, named specifically - the applications, hosts or ranges, and the roles tested. "The production environment" is not a scope.
  • What was excluded, and why. An undocumented exclusion is indistinguishable from a gap in coverage.
  • The methodology followed, by name and version - OWASP WSTG, ASVS, the API Top 10, or your provider’s documented equivalent. PCI requires this explicitly.
  • Whether testing was authenticated, and against which roles. An unauthenticated test of an authenticated application covers very little of it.

Establishing what was found

  • Each finding with a severity and the basis for it - a CVSS vector, so the rating can be checked rather than taken on trust.
  • Evidence for each finding: the request and response, or the steps that reproduce it. This is also what lets your engineers fix it without a meeting.
  • The affected component per finding, so remediation can be assigned and the fix verified against something specific.
  • Remediation guidance that names what to change, not a link to a general article.
  • An explicit statement of limitations - what a point-in-time test does not establish.

If it helps to see all of that in one place rather than as a list, there is a how we test on this site with every section above in it, including the compliance mapping variants. Nothing is gated behind an email.

The retest is the part people fail

A report showing critical findings, with nothing after it, is evidence that you have a problem. That is the most common way an otherwise good engagement fails to close an audit item.

PCI DSS states it outright at 11.4.4: exploitable vulnerabilities and security weaknesses found during penetration testing are corrected, and testing is repeated to verify the corrections. The other three frameworks do not use that language, but every assessor asks the same question in some form, because a control that identifies vulnerabilities and does nothing about them is not operating.

  • Get the retest in writing, as a dated verdict per finding rather than a sentence saying issues were addressed.
  • Keep the trail: when it was reported, when it was fixed, who verified it and when. A dated record answers the follow-up question before it is asked.
  • Where something is not going to be fixed, record the risk acceptance properly - who accepted it, on what basis, and when it will be revisited. An accepted risk is a defensible answer; an ignored finding is not.
  • Check whether your provider charges for the retest. If it is billed as new work, the cost of being audit-ready is higher than the quote you compared.

How we handle this

Retests are included in the engagement rather than billed separately. You mark a finding fixed and request a retest from the same screen; a tester verifies it by hand and records whether it passed or failed, and every transition is written to an append-only trail with the actor and the time. That record is the evidence, and it is exportable with the report.

Assembling the evidence pack

What you hand over is usually a small bundle rather than a single document. Assembling it before it is asked for turns a week of back-and-forth into one email.

  • The report itself, in a format the recipient can actually open and quote from. Word or PDF for an assessor, a spreadsheet where they want to track findings alongside their own workpapers.
  • The retest record, or the report reissued with verified fixes marked.
  • The scope and rules of engagement, or the section of the report that reproduces them.
  • Evidence of the cadence: the previous report, if the framework expects testing at an interval.
  • Your remediation record for anything still open, including risk acceptances with dates and owners.
  • Where a framework mapping is expected, the report grouped under that framework’s own requirement areas rather than a mapping written by hand afterwards.

On that last point: the same engagement can be exported against PCI DSS, SOC 2, ISO 27001 or HIPAA with the findings grouped under that framework’s requirement areas, so the mapping is part of the deliverable rather than a spreadsheet somebody maintains.

The three ways this goes wrong late

The scope did not match the boundary

A test of the marketing site when the assessor cares about the payment path. A test of one application when the system description names three. This is discovered when the evidence is reviewed, which is the worst possible time, and the only fix is another engagement.

There was no time left for remediation

The test finished a week before the deadline, found two criticals, and there is now a report proving an unresolved problem and no time to resolve it. Book so that the retest lands before the date, not the test.

The report cannot support a coverage claim

No dated window, no named scope, no methodology, no evidence per finding - often the output of an automated scan presented as a penetration test. It may be perfectly accurate and still be unusable, because an assessor cannot establish from it what was examined. If you are comparing providers, ask each for a sample report and check it against the list in this guide before you compare prices.

FAQ

Questions this raises

Does SOC 2 actually require a penetration test?

No. The Trust Services Criteria do not name penetration testing as a required control. They require that vulnerabilities are identified, evaluated and remediated, and that controls are monitored, which is CC4 and CC7 territory. A recent test report is the evidence most organisations offer for those criteria and most auditors expect it, but that is convention rather than a written requirement. If your auditor has told you it is required, ask which criterion they are mapping it to - the answer is useful for scoping the test.

How often do we need to test?

PCI DSS is the only one of the four that sets an interval: at least once every 12 months and after any significant change, for both internal and external testing. SOC 2, ISO 27001 and HIPAA set no interval, and annual testing is convention rather than a requirement. In practice the trigger that matters more than the calendar is significant change - a re-platform, a new authentication provider, a new payment path.

Will an automated scan do instead?

Sometimes, for some purposes, and it is worth knowing which. PCI distinguishes vulnerability scanning at Requirement 11.3 from penetration testing at 11.4 and requires both, so a scan does not substitute there. Where a framework asks for outcomes rather than a named activity, an assessor may accept scanning evidence for part of it. What a scan cannot evidence is anything about authorisation or business logic, because it cannot test them - which is also why those are the findings that matter most in this guide.

Our auditor asked for the previous year’s report as well. Is that normal?

Yes, particularly where a framework expects testing at an interval or where the examination period spans two tests. It is worth keeping reports and retest records for at least two cycles for that reason, along with the remediation record for anything that was still open at the time.

Can we show our auditor a report with unresolved findings?

Yes, and it is normal - what matters is what accompanies it. A finding with a dated remediation plan and an owner, or a properly recorded risk acceptance, is a control operating. The same finding with nothing beside it is the thing an assessor raises. The pattern to avoid is a report handed over with no record of what happened next.