Vulnerability scan vs penetration test: what each one finds
A vulnerability scan and a penetration test both end in a list of security problems, which is where the confusion starts. One is software checking for known weaknesses. The other is a person proving what an attacker could actually do. Here is what each one finds, where automated penetration testing fits between them, and what the standards expect.
Key takeaways
- A vulnerability scan checks your systems against a library of known weaknesses. A penetration test is a person proving what an attacker could actually do. They answer different questions, and neither substitutes for the other.
- Scanners are fast, consistent and cheap to repeat. For published CVEs, missing patches, default credentials, weak TLS and exposed services they are the right tool, and they should run often.
- A clean scan means none of its checks fired against what it could reach. It says nothing about authorisation, business logic or how small findings chain together.
- Automated penetration testing is strongest on internal networks, where the attack techniques against standard components are well documented. It does not know your authorisation rules or what your application is for.
- PCI DSS v4.0.1 treats them as separate requirements: scanning under 11.3 at least once every three months, penetration testing under 11.4 at least once every 12 months, and both after any significant change.
The difference, in plain terms
A vulnerability scan is automated. Software checks your hosts, services and applications against a library of known weaknesses and reports every check that matched: a missing patch, a published CVE, a default password, a weak TLS configuration. It is fast, consistent and cheap enough to run every week.
A penetration test is a person, using tools for coverage, trying to do what an attacker would and proving what is possible. The output is a set of confirmed findings with the evidence for each, including flaws no scanner detects: one customer reading data that belongs to another, or two harmless-looking weaknesses chained into an account takeover.
Put shortly, a scanner finds what is there and a tester finds what is broken, and neither answer stands in for the other.
| Question | Vulnerability scan | Penetration test |
|---|---|---|
| Who does the work | Software, scheduled by a person. | A tester, with tools for coverage and reasoning by hand. |
| What it looks for | Known weaknesses: missing patches, published CVEs, default credentials, weak TLS, exposed services. | Anything that lets somebody do what they should not, including flaws with no signature. |
| How a finding is established | A response, banner or version matches a known check. | The issue is exploited, or proven far enough to be certain, and the evidence is recorded. |
| Authorisation and business logic | Not assessed. The tool does not know who should be allowed to do what. | Central to the work, tested across roles, accounts and tenants. |
| Chaining | Each result is scored on its own. | Weaknesses are combined to show what they lead to. |
| False positives | Expected, so results need triage before anybody acts. | Far fewer, because each finding is confirmed before it is reported. |
| PCI DSS v4.0.1 | Requirement 11.3: at least once every three months and after any significant change. | Requirement 11.4: at least once every 12 months and after any significant change. |
What a vulnerability scanner does well
Scanners deserve a fair hearing. A scanner checks every host in range the same way every time and costs little to run again tomorrow, and for the problems it is built for it beats a person: nobody should pay a tester to compare version strings by hand.
- Published vulnerabilities in software it can identify, and the patches that fix them.
- Default credentials on devices, panels and services.
- Weak TLS, expired certificates and missing security headers.
- Exposed services and administrative interfaces, and operating systems past end of support.
- Change over time: a port that opened or a host that appeared since last week, the drift a periodic test misses.
Web application scanners also send crafted input to parameters, so they find some injection flaws and reflected cross-site scripting in your own code. What they recognise is still a pattern.
Unauthenticated and authenticated scans
An unauthenticated scan sees what anybody on the network sees: open ports, service banners and responses to probes. Because it infers so much from version information, it can be wrong in both directions. A server whose vendor backported a security fix without changing the advertised version is flagged for a flaw it no longer has, and a service that hides its version gets a pass it has not earned.
An authenticated scan logs in with credentials you provide and reads installed packages, patch levels and configuration directly, which removes most of that guesswork. PCI DSS now requires it for the quarterly internal scans under requirement 11.3.1.2, mandatory since 31 March 2025. For a web application, authenticated means the scanner can crawl behind the login, but it still sees the application as one user.
Where scanning stops
Every limit of a scanner follows from one fact: it can only report what it has a check for, against what it can reach.
- Known patterns only. A flaw in your own code that matches nothing in the tool library produces no result, however serious.
- False positives. Results inferred from banners and versions need confirming before engineers spend a day on them.
- False negatives. Pages the crawler never reached, multi-step workflows, ports outside the scanned range and requests a firewall dropped produce no finding, and no visible gap either.
- No grasp of authorisation. A scanner that requests an invoice and receives one cannot tell whether it was entitled to it. Some tools can compare two accounts and flag responses that differ, but deciding whether a difference is a breach requires knowing the rules.
- No grasp of business logic. A price accepted from the browser or a replayable refund involves only well-formed requests and successful responses, so there is nothing to match.
- No chaining. Results are scored in isolation, so two findings that combine into something serious appear as two unremarkable lines.
A clean scan is not a clean bill of health
No high or critical results means that none of the checks the tool ran fired against what it could reach. It does not mean nothing exploitable is there. Broken access control, a price taken from the browser and two low-severity issues that combine into an account takeover produce no signature, so they produce no result. Read a clean scan as evidence that the known-weakness layer is under control, which is worth having, and as nothing more.
What a manual penetration test adds
A manual penetration test starts where the scan stops. The tester runs tooling for coverage, then works out what the system is for and tries to make it do something it should not. Four things come out of that which no scan produces.
- Proof. Each finding is exploited, or taken far enough to be certain where going further would cause damage, and the request and response that demonstrate it are recorded. That is why a test produces far fewer false positives than a scan.
- Authorisation testing. The tester signs in as each role, ideally with two accounts per role and, in a multi-tenant product, two tenants, and tries every object against every identity. This is where the most serious application findings live, and what a penetration test actually finds goes through them class by class.
- Business logic. A tester who knows that an order should be paid for before it ships can check whether the application agrees.
- Chaining. An API that lists the identifiers of other customers is a minor disclosure, and an endpoint that accepts any identifier without checking ownership is a finding of its own. Together they are a bulk export of every customer record. A tester reports the chain, because the chain is the risk.
Manual penetration testing does not mean testing without tools, and no honest test is entirely manual. What matters is which half does which job, and the split is set out on how we test.
Automated penetration testing: what it does well and where it stops
Automated penetration testing is the name for tools that go a step past scanning. Rather than reporting that a weakness might exist, they try to use it, then try the next step from wherever that left them. Automated pen testing is genuinely good at part of this work, so it is worth being precise about which part.
Where automation is strong
An internal Windows network is assembled from the same components almost everywhere: a directory service, file shares, authentication protocols, group memberships and service accounts. The ways an attacker moves through one are well documented and repeat between organisations: a password reused across machines, credentials readable on a share, a service account whose password can be cracked offline, group rights nobody reviewed.
Because the techniques are known and the components are standard, they can be encoded. A tool can map the relationships between accounts, machines and groups, walk the paths from an ordinary user to domain administrator, and repeat the exercise every week. That answers two questions well: does a path to full control exist today, and did the fix made last month close it? Run between manual engagements, it catches drift that an annual network penetration test would see only once a year.
Where it stops
Applications are not assembled from standard parts. The flaws that matter in a web application or an API are specific to your code: which role may read which record, what a discount may do, which step must come before which. A tool can be given two accounts and told to compare what each one sees, but not easily the rule that says which differences are correct, and without that rule, data belonging to somebody else looks exactly like your own.
Unattended exploitation also has to be cautious. A person can prove a flaw far enough to be certain and stop. A tool with nobody watching either avoids anything that might change data, which rules out much of what an application does, or takes actions nobody reviews. It also follows only the paths it was written to follow. Newer tools, including some built on language models, attempt more of the application layer. Judge them as you would a tester: by what was proven, against which roles, and what was not covered.
The PCI Security Standards Council describes the difference in its penetration testing guidance: penetration testing is essentially a manual process, in which tools aid the tester and relieve some of the repetitive work, and judgement is needed to identify attack vectors that automated means typically cannot. That document is supplementary guidance rather than a requirement, and it predates the current numbering.
How scanning and testing fit together
The useful question is not scan or test, but which question each one answers and how often you need the answer. A scan is cheap enough to run weekly; a test is run periodically and after significant change. Run scanning first and testing second, which is the order we work in, and each improves the other. Scanning clears the known-weakness layer, so testing days go to authorisation and logic rather than missing patches. Between tests, scanning notices when the ground moves, so the next test starts from the estate as it is now.
Our own monitoring between engagements is built on that division. On the plans that include it, each asset, once you have proved you own it, gets a lightweight check every 24 hours: the hostname is resolved, the TLS certificate is read and open ports are compared with the previous snapshot. A full scan runs weekly or monthly as you choose, deliberately not daily, because a deep scan of an unchanged target returns the same answer it returned yesterday. Host discovery runs weekly from certificate transparency records. Monitoring supplements manual testing and does not replace it, and anything a scheduled scan raises is kept separate from the findings a tester wrote.
When to use which
A scan is the right tool when
- You need to know, on a schedule, whether known weaknesses are present across everything you expose.
- You have patched, reconfigured or changed infrastructure and want to confirm what actually changed.
- A test is coming, and you would rather the tester did not spend days reporting missing patches.
A penetration test is the right tool when
- The risk sits behind a login: user roles, tenants, payments, anything that moves money or changes state.
- You are launching, or have significantly changed authentication, authorisation or a payment flow.
- A customer, an auditor or an investor needs assurance. They are asking what a person found.
- A framework names it, as PCI DSS does, or a year has passed. Annual testing plus testing after significant change is the common baseline, though only some frameworks write it down.
If anything sensitive sits behind a login you will usually want both, and a vulnerability assessment and penetration test combines them in one engagement. The scoping guide covers what drives the effort.
What compliance frameworks expect
PCI DSS settles the question in writing. Version 4.0.1, current since version 4.0 was retired at the end of 2024, places vulnerability scanning under requirement 11.3 and penetration testing under requirement 11.4, as separate obligations with separate evidence.
| Requirement | What it asks for | How often |
|---|---|---|
| 11.3.1 and 11.3.1.3 | Internal vulnerability scans by qualified, organisationally independent personnel, with high-risk and critical vulnerabilities resolved and confirmed by rescanning. | At least once every three months, and after any significant change. |
| 11.3.1.2 | Internal scans performed as authenticated scans, with systems that cannot accept credentials documented. | Applies to the quarterly internal scans. A best practice until 31 March 2025, required since. |
| 11.3.2 and 11.3.2.1 | External vulnerability scans. The quarterly scans must be run by a PCI SSC Approved Scanning Vendor (ASV) and meet the council criteria for a passing scan. Scans after a change can be run by qualified, independent personnel. | At least once every three months, and after any significant change. |
| 11.4.2 and 11.4.3 | Internal and external penetration testing to a documented methodology, by a qualified, organisationally independent resource. | At least once every 12 months, and after any significant infrastructure or application upgrade or change. |
A passing ASV scan answers 11.3.2 and nothing in 11.4, and a penetration test answers 11.4 and nothing in 11.3. Which parts apply to you depends on how you validate, since the shorter self-assessment questionnaires include only part of requirement 11. The PCI DSS penetration testing guide covers 11.4 in detail, including segmentation testing and retesting.
SOC 2 and ISO 27001 name neither activity as a requirement. Under SOC 2, criterion CC7.1 on detecting vulnerabilities over time is where scanning evidence usually lands, with a test presented beside it. ISO 27001 asks under Annex A control 8.8 for technical vulnerabilities to be identified, evaluated and addressed, and leaves the method to you. Under both, regular scanning plus an annual test is common practice rather than a written rule.
Whatever the framework, neither a scan nor a test is a certification. Each is evidence that may support an assessment, and the assessor decides what satisfies the requirement.
Keep reading
- What a penetration test actually findsThe classes of flaw a manual test turns up in applications and APIs, and why scanners miss most of them.
- PCI DSS penetration testing requirementsRequirement 11.4 sub-requirement by sub-requirement, including segmentation testing and the obligation to retest.
- Attack surface monitoringA daily check, a full scan weekly or monthly, and weekly host discovery between engagements.
- Vulnerability assessment and penetration testingScanning for coverage, then manual exploitation of what matters, in one engagement.