Black box, grey box and white box penetration testing
The three labels describe one decision: how much the tester is told before testing starts. That decides where the testing time goes and what the report can honestly say was examined, which matters far more than which option sounds most like a real attack. Here is what each one involves, what it finds and misses, and how to write the choice into a scope.
Key takeaways
- The labels describe one thing: how much the tester is given before testing starts. Black box gets a target, grey box adds accounts and documentation, white box adds the source code. None of them is a grade of thoroughness.
- Black box is realistic about an outsider with no foothold, and pays for it in discovery time. Choose it when that outsider is the question, or when the exercise is really about detection and response.
- Grey box is the usual default for application tests because authorisation is tested as a matrix of roles and tenants, and that needs a working account on each side of every boundary.
- White box is worth paying for where the risk sits in logic rather than exposure. It overlaps with secure code review without replacing it, and the code reviewed has to match the code deployed.
- Write the inputs into the scope, not the colour. The standards draw the lines in different places, so name the accounts, the documents, the repositories and who is told.
What is the difference?
The difference between black box, grey box and white box penetration testing is how much the tester is given before testing starts. It is a scale of disclosure, from told nothing to told everything, and not a scale of quality. What moves along it is where the testing time goes, and what the report can honestly say was examined.
- In a black box test the tester receives the target and nothing else: a hostname, an IP range, an app store listing. No accounts, no documentation, no architecture overview, so anything they learn, they learn the way an outsider would.
- Grey box penetration testing (gray box penetration testing in US spelling) gives partial knowledge. For an application that normally means a working account for each user role, an architecture overview and whatever API documentation exists, but not the source code.
- In a white box test the tester also gets the source code, and usually the infrastructure configuration and design documents, so they can read how the system is meant to work as well as watch how it behaves.
These definitions follow the PCI Security Standards Council penetration testing guidance, which separates the three by whether the client provides no information, partial details or full details. Not every document draws the lines in the same place, as the last section shows.
The three approaches side by side
| What changes | Black box | Grey box | White box |
|---|---|---|---|
| Accounts and roles | None. Any account is obtained as an attacker would obtain it. | One per role, ideally two, across two tenants if multi-tenant. | As grey box. |
| Documentation | The target name only. | Architecture overview, API documentation, and what each role may do. | As grey box, plus design documents. |
| Source code and configuration | None. | None. | Read access to the relevant repositories, often with infrastructure configuration. |
| Where the time goes | Reconnaissance and mapping before anything behind a login is reached. | The authorisation model, the workflows and the API behind the interface. | Reading the riskiest code paths, then confirming against the running system. |
| Finds well | Exposed hosts and services, unauthenticated endpoints, weak sign-up, login and reset flows. | Access between roles and tenants, business logic, flaws in authenticated functions. | Missing checks, flawed logic, race conditions, cryptographic misuse, secrets in code. |
| Tends to miss | Most of what sits behind a login, and every role it cannot obtain. | Code paths the interface rarely exercises, and how far an outsider gets unaided. | Little, if the running system is still tested. The usual gap is code that differs from what is deployed. |
| Best suited to | The outsider question, and exercises aimed at detection and response. | Most application and API tests, and anything with more than one role. | High-risk logic, bespoke cryptography, critical applications before launch. |
Effort follows from the inputs. Black box spends the early window on discovery the client could have shortcut, and leaves untested whatever it cannot reach. Grey box puts that time into the authorisation model, and because the roles and endpoints are known, the report can state coverage against them. White box adds reading time and buys depth with it. None is cheaper by default: each changes what a budget buys, as the cost guide sets out.
Black box: the outsider with no foothold
Black box penetration testing answers one question well: what can somebody with no account, no inside knowledge and no help reach from where they stand? The answer covers forgotten hosts, administrative panels meant to be internal, services that should never have been public, and whatever sign-up, login and password reset give away to a stranger.
The cost is time. Before anything else, a black box tester has to map the application, guess at its roles, find the API and try to obtain an account, and an attacker can take as long as they like over that. A test has a fixed window. Time spent rediscovering what you already know is not spent behind the login, and without open registration the test may never get behind it, so the report can only describe what was reachable, not what was examined.
Black box is the right call in fewer situations than its reputation suggests:
- The outsider really is the question, such as an external perimeter where the concern is what a stranger can reach.
- The exercise is about whether your team detects and responds, which is a different engagement, covered below.
- It is the first phase of a staged test.
If the question is only what you have exposed, scanning answers it more cheaply than manual testing time.
Black box and covert testing are different decisions
Black box describes what the tester knows. Whether your own staff know a test is running is a separate decision, which NIST SP 800-115 calls overt or covert testing (section 2.4.2). Covert testing runs without the knowledge of the IT staff but with the permission of upper management, and NIST describes its purpose as examining the damage an adversary could cause rather than identifying vulnerabilities. It is often slow and costly because of its stealth requirements. If the objective is to learn whether anyone would notice, that is a red team engagement, and it can start from a valid account as easily as from nothing.
Realistic is not the same as thorough
The usual argument for black box is that it is the most realistic option. It is realistic about one thing: the starting position of an outsider who has been told nothing. It is not realistic about time, because an attacker has no fixed window, and often not about the attacker. For a product anyone can sign up to, or one with many customers, the attacker with the best odds already has a login. A black box result partly measures how hard your system is to understand from outside, and that protection lasts only until somebody spends the time.
Grey box: the default for application tests
Grey box is the usual default for application testing, for a structural reason. The findings that matter most in an application are mostly about access: a customer reading an order that belongs to somebody else, a standard user reaching an administrative function, one tenant seeing into another. Broken access control cannot be tested from one side of a boundary. It needs an identity on each side.
Authorisation is tested as a matrix, each role tried against every function and object the others can reach. A black box tester can only fill in the rows for roles they can obtain, which for most products means one self-registered user, or nobody. Grey box hands over the rest: an account per role, two where access between peers matters, and two tenants where the product is multi-tenant.
The documentation matters as much as the accounts, because a tester can only call something a finding if they know it was not meant to happen. Whether a support agent may export a customer list is a business decision, and a note on each role settles it faster than inference. API documentation is a starting map: the tester still enumerates what is actually exposed, because undocumented endpoints are where documentation stops being true.
Grey box as the default is common practice, not a written rule, but the published guidance points the same way. The OWASP Web Security Testing Guide says penetration testing is commonly known as black box testing, and adds in the same passage that in many cases the tester is given one or more valid accounts. The PCI SSC guidance says PCI DSS penetration tests are typically performed as white box or grey box assessments, and that these give more accurate results and a more comprehensive test than a pure black box assessment. NIST SP 800-115 describes internal network testing the same way, with assessors often given access as general users (section 2.4.1), which is an assumed-breach network test by another name.
It is also what we recommend for most web application and API tests, and the questions we ask before quoting include a request for one account per role, because a role nobody can sign into is a role nobody can test.
Grey box has limits. It does not measure how far an outsider gets unaided, though sign-up, login and reset are still tested. And it sees code only through behaviour, so a branch the interface never exercises, or a race that needs one exact sequence of requests, can go unnoticed. That is the gap white box closes.
White box: when reading the code beats guessing
White box is worth paying for where the risk sits in logic: an authorisation model with many branches, a payment flow where two requests can race, bespoke cryptography, or a fix that has to be confirmed everywhere a pattern occurs.
The OWASP guide says many serious vulnerabilities cannot be detected by any other form of analysis or testing, and names concurrency problems, flawed business logic, access control problems and cryptographic weaknesses as issues source code review is particularly suited to finding. Among its listed limits: the code deployed may differ from the code reviewed, and runtime errors are hard to see in source. NIST SP 800-115 says that for custom applications, analysing the source tends to be more efficient and cost-effective than testing without it, but cannot detect defects in the interfaces between components or problems introduced during compilation, linking or installation-time configuration. It recommends black box techniques for interactions between components and with users and other systems (Appendix C).
So a white box penetration test still tests the running system: the code tells the tester where to look and explains what they find. It is not automatically longer, since reading code replaces guessing. The added time depends mostly on how much code is in scope, and whether all of it is read or only the paths that carry the risk. It is not the default mainly because not every organisation can share source with a third party, and for many products the accounts alone reach the access control risk that matters most.
White box testing and secure code review
The two overlap, and they are different purchases. In a white box test the running system is the target and the code is a guide to it. In a secure code review the code is the target: the reviewer works through the risky paths, such as authorisation, cryptography and input handling, and reports defects at the file and line whether or not anything can reach them yet. A review finds the ownership check present on three handlers and missing from the fourth. A test proves which gaps an attacker can actually use, which is why the two are often bought together.
If you share code, read-only access to the relevant repositories is enough. Name the branch or commit in the scope and confirm that it matches the environment being tested, because a review of the wrong commit is a review of software you do not run.
Authenticated testing and multiple roles
Box colour and authenticated testing are separate questions that often get treated as one. Authenticated testing means working with valid sessions, and grey box and white box tests almost always include it. A black box test of a product with open sign-up becomes authenticated once the tester registers, but only as whatever role sign-up hands out.
Most business applications have roles a stranger cannot sign up for: administrators, support staff, finance approvers, partners with their own API keys. A role without an account is a role nobody tested, so if the report is to say anything about the boundaries between them, the accounts have to be provided.
The credentials section of the scoping guide has the detail: an account per role and two where peers matter, two tenants with data that is easy to tell apart, and multi-factor authentication settled before the start date.
Writing the choice into a scope
Writing grey box into a statement of work settles less than it seems to, because even the standards use the labels inconsistently. The PCI SSC guidance defines them by how much the client hands over. NIST SP 800-115 uses white box for analysing source code, black box for testing without it, and grey box for the combination. The OWASP guide describes its own method as based on the black box approach, with a tester who knows little or nothing about the application, yet expects accounts in many cases. The same engagement can be white box in one document and grey box in another.
So write down the inputs, not the colour:
- The question the test is meant to answer, in one sentence.
- Accounts: which roles, how many of each, in which tenants, and who provisions them by when.
- Whether the tester may register their own accounts, and whether those are in scope.
- Documentation: architecture overview, API specification, and what each role may do.
- Source code, if any: which repositories, and which branch or commit.
- Whether the web application firewall and rate limiting are relaxed for the testing source addresses.
- Whether your monitoring team knows the test is running, recorded as its own decision.
- Any staging, such as an unauthenticated phase before the accounts are handed over.
A staged test suits an engagement where the outsider view matters but should not take the whole budget. Run the uninformed part first, because a tester cannot unlearn what they have been told. NIST SP 800-115 says the same of network testing: where internal and external testing are both done, the external part usually comes first, so the assessors do not carry insider knowledge into it (section 2.4.1).
The report should then repeat what was provided, so a reader can tell what the test could and could not have seen. That belongs in the scope statement, the part of a report that outlives the findings. If it is still unclear which approach answers your question, that is what a scoping call is for, and we will say which one we would recommend and why.
Keep reading
- How to scope a penetration testWhat a tester needs before quoting, the test accounts to prepare, and the scoping mistakes that waste an engagement.
- What a penetration test costsThe factors that move the number, including how much access the tester is given.
- Secure code reviewManual review of authorisation, cryptography and injection paths, with each finding at the file and line.
- Types of penetration testingWhat each type examines, from external networks to APIs, and which of them a scope needs.