The audit evidence checklist for a penetration test
Most people buy a penetration test because somebody asked them for one - an auditor, a customer, a QSA. This is the checklist for making sure what you buy is what they will accept: what each framework actually requires, what belongs in the evidence pack, and the three places this goes wrong late enough to move an audit date.
Key takeaways
- Only PCI DSS writes the requirement down. SOC 2 and ISO 27001 do not mandate a penetration test at all - auditors ask for one because it is the most convenient evidence for criteria written in terms of outcomes.
- The single most common gap is not the test. It is the absence of evidence that findings were fixed and verified, which PCI DSS requires explicitly at 11.4.4 and every other framework expects in practice.
- A report with no dated testing window, no named scope and no methodology is hard for an assessor to accept whatever it found, because they cannot establish what it covered.
- Scope has to match the boundary the framework cares about - the cardholder data environment, the system described in your SOC 2 description, the ISMS scope. A test of the wrong boundary is a real test and useless evidence.
- Book so the retest lands before your evidence deadline, not the test. Remediation time is the part people forget to budget.
What each framework actually requires
Start here, because the four are not equivalent and treating them as one requirement is how people over-buy or under-deliver. One of them names penetration testing and sets a cadence. The other three do not mention it as a mandatory control at all, and ask for outcomes that a penetration test happens to be a convenient way of evidencing.
| Framework | What the text requires | What is usually asked for |
|---|---|---|
| PCI DSS v4.0 | Requirement 11.4 requires a defined methodology (11.4.1), internal and external penetration testing at least once every 12 months and after significant change (11.4.2, 11.4.3), and that exploitable vulnerabilities found are corrected and testing repeated to verify the correction (11.4.4). | The report, the methodology it followed, and evidence of the repeat testing at 11.4.4. This is the one framework where the retest is written into the requirement rather than being good practice. |
| SOC 2 | The Trust Services Criteria do not require a penetration test. They require that vulnerabilities are identified, evaluated and remediated, and that controls are monitored - CC4 and CC7 in the common criteria. | A recent test report is the usual evidence offered for those criteria, together with what you did about the findings. An auditor may accept other evidence; most ask for this because it is the cleanest. |
| ISO/IEC 27001:2022 | Annex A.8.8 requires that technical vulnerabilities are identified, evaluated and addressed. Penetration testing is not named as a mandatory control anywhere in the standard. | A test report as evidence that A.8.8 is operating, plus the record of what was done with the findings. Certification bodies vary in how much they want to see. |
| HIPAA Security Rule | Section 164.308(a)(8) requires periodic technical and non-technical evaluation. Section 164.308(a)(1)(ii)(A) requires a risk analysis. Neither names penetration testing. | Technical testing evidence supporting the evaluation, and a demonstrable link from what was found to your risk analysis and remediation. |
The sentence to keep in mind
A penetration test is evidence that may support an assessment. It is not a certification or an attestation, and any provider who tells you their report certifies you against a standard is either being loose with language or does not understand the standard. Your assessor decides what satisfies it.
Before you book: what to settle
Every item here is cheap to answer now and expensive to answer late. The last two are the ones that move dates.
- Who is asking, and for what. An auditor mid-examination, a customer security questionnaire and a QSA all want different things from the same engagement. Ask for the request in writing and read what it actually says.
- Which boundary has to be covered. The cardholder data environment, the system described in your SOC 2 system description, the scope of your ISMS. A test of the wrong boundary is a real test and useless as evidence.
- What counts as a significant change since the last one. PCI names this explicitly; if you have re-platformed, added an authentication provider or moved a component, the previous test may not carry.
- Whether internal as well as external testing is required. PCI requires both. The others do not distinguish, but an auditor looking at a system with an internal trust boundary often asks.
- The date the evidence is needed, not the date the test can start. Then work backwards through remediation and retest, which is where the real time goes.
- Who will own the findings internally. A report delivered to somebody with no authority to schedule engineering work is how a finding is still open at the next audit.
Budget for remediation, not just the test
A five to ten day engagement is the short part. Fixing what it finds and having the fix verified is what determines whether you have evidence by your date. If the deadline is six weeks away, book now rather than in a month.
What the report has to contain
Assessors differ, but the items below are what makes a report usable as evidence rather than merely interesting. The first group exists so a reader can establish what was examined; without them a report cannot support a claim about coverage however good the testing was.
Establishing what was covered
- A dated testing window, not just a publication date. An assessor needs to place the work inside the period under examination.
- The scope, named specifically - the applications, hosts or ranges, and the roles tested. "The production environment" is not a scope.
- What was excluded, and why. An undocumented exclusion is indistinguishable from a gap in coverage.
- The methodology followed, by name and version - OWASP WSTG, ASVS, the API Top 10, or your provider’s documented equivalent. PCI requires this explicitly.
- Whether testing was authenticated, and against which roles. An unauthenticated test of an authenticated application covers very little of it.
Establishing what was found
- Each finding with a severity and the basis for it - a CVSS vector, so the rating can be checked rather than taken on trust.
- Evidence for each finding: the request and response, or the steps that reproduce it. This is also what lets your engineers fix it without a meeting.
- The affected component per finding, so remediation can be assigned and the fix verified against something specific.
- Remediation guidance that names what to change, not a link to a general article.
- An explicit statement of limitations - what a point-in-time test does not establish.
If it helps to see all of that in one place rather than as a list, there is a how we test on this site with every section above in it, including the compliance mapping variants. Nothing is gated behind an email.
The retest is the part people fail
A report showing critical findings, with nothing after it, is evidence that you have a problem. That is the most common way an otherwise good engagement fails to close an audit item.
PCI DSS states it outright at 11.4.4: exploitable vulnerabilities and security weaknesses found during penetration testing are corrected, and testing is repeated to verify the corrections. The other three frameworks do not use that language, but every assessor asks the same question in some form, because a control that identifies vulnerabilities and does nothing about them is not operating.
- Get the retest in writing, as a dated verdict per finding rather than a sentence saying issues were addressed.
- Keep the trail: when it was reported, when it was fixed, who verified it and when. A dated record answers the follow-up question before it is asked.
- Where something is not going to be fixed, record the risk acceptance properly - who accepted it, on what basis, and when it will be revisited. An accepted risk is a defensible answer; an ignored finding is not.
- Check whether your provider charges for the retest. If it is billed as new work, the cost of being audit-ready is higher than the quote you compared.
How we handle this
Retests are included in the engagement rather than billed separately. You mark a finding fixed and request a retest from the same screen; a tester verifies it by hand and records whether it passed or failed, and every transition is written to an append-only trail with the actor and the time. That record is the evidence, and it is exportable with the report.
Assembling the evidence pack
What you hand over is usually a small bundle rather than a single document. Assembling it before it is asked for turns a week of back-and-forth into one email.
- The report itself, in a format the recipient can actually open and quote from. Word or PDF for an assessor, a spreadsheet where they want to track findings alongside their own workpapers.
- The retest record, or the report reissued with verified fixes marked.
- The scope and rules of engagement, or the section of the report that reproduces them.
- Evidence of the cadence: the previous report, if the framework expects testing at an interval.
- Your remediation record for anything still open, including risk acceptances with dates and owners.
- Where a framework mapping is expected, the report grouped under that framework’s own requirement areas rather than a mapping written by hand afterwards.
On that last point: the same engagement can be exported against PCI DSS, SOC 2, ISO 27001 or HIPAA with the findings grouped under that framework’s requirement areas, so the mapping is part of the deliverable rather than a spreadsheet somebody maintains.
The three ways this goes wrong late
The scope did not match the boundary
A test of the marketing site when the assessor cares about the payment path. A test of one application when the system description names three. This is discovered when the evidence is reviewed, which is the worst possible time, and the only fix is another engagement.
There was no time left for remediation
The test finished a week before the deadline, found two criticals, and there is now a report proving an unresolved problem and no time to resolve it. Book so that the retest lands before the date, not the test.
The report cannot support a coverage claim
No dated window, no named scope, no methodology, no evidence per finding - often the output of an automated scan presented as a penetration test. It may be perfectly accurate and still be unusable, because an assessor cannot establish from it what was examined. If you are comparing providers, ask each for a sample report and check it against the list in this guide before you compare prices.
Keep reading
- A full worked example reportEvery section this checklist asks for, in a real report you can read on the page.
- Penetration testing for SOC 2Which criteria a test actually evidences, and what an examination looks for.
- PCI DSS penetration testing requirementsRequirement 11.4 read closely, including the repeat testing at 11.4.4.
- Compliance reporting in VexilThe same engagement exported against four frameworks, with control mapping.