GetPentest

Types of penetration test and when to use each

The type of test you buy decides what the report can tell you. Most companies buy black box testing because it sounds more like a real attack, and most of them get less for their money because of it.

Last reviewed 2026-08-16Written by Jacob Masse, TrazTech Inc.

Penetration tests are described two ways at once: by how much information the tester is given, and by what the engagement is trying to achieve. The first gives you black box, grey box and white box. The second separates a penetration test from a red team exercise. Choosing wrongly on either axis wastes a good share of the budget.

Six of these have enough scoping detail to need a page of their own: wireless, Active Directory, Kubernetes and containers, CI/CD pipelines, thick clients and IoT and embedded devices.

The four in one table

Penetration test types compared
Type Tester is given Best for Relative cost
Black box A name or an IP range and nothing else Testing what is discoverable from outside, and your exposure surface Baseline
Grey box Credentials, roles, architecture notes, sometimes an API specification Almost every application and internal network engagement Baseline to 1.5x
White box Source code, configuration, design documents, full access Dense business logic, cryptography, anything safety relevant 1.5x to 2.5x
Red team An objective, and permission to reach it any way they can Mature teams testing detection and response, not vulnerabilities 4x and up

In one line each.

Black box
The tester starts where a stranger starts, with a name or an address. Measures your exposed surface, and spends part of your budget on discovery you could have handed over.
Grey box
The tester starts where a customer or an employee starts, with credentials and enough architecture to be useful. The correct default for almost every application and internal engagement.
White box
The tester starts with the source. Buys depth in a narrow area, which is what you want for cryptography, payment flows and anything where a mistake ends the company.
Red team
The tester starts with an objective rather than a scope, and nobody internally knows. Measures detection and response rather than vulnerabilities, and is set out on red team assessment.

Black box

The tester starts with what a stranger on the internet has: your domain name, or a range of addresses. Everything else is discovered. It is the closest simulation of an opportunistic external attacker, and it answers one question well: what an outsider can find and reach.

When it is the wrong choice. For any application with a login. You are paying senior day rates for reconnaissance you could have handed over in an email, and reconnaissance eats a third or more of a short engagement. A real attacker has unlimited time and your tester has five days, so the comparison to a real attack flatters rather than describes. You get a report about your perimeter and TLS configuration, and silence about the authorization model where the risk lives.

Buy black box when the exposure surface itself is the question: after a merger, after a cloud migration, or when you do not know what your company has published to the internet.

Comparing firms for this? Tell us what you need and it goes to the ones in the directory that do this work. No charge, and no phone number required.

Grey box

The tester gets credentials for each role, an architecture overview, and usually an API specification. It skips the discovery phase and spends the budget on the testing that finds real problems, particularly the cross-role and cross-tenant work described on web application security testing.

A real attacker gets credentials. In a SaaS product they sign up for a free trial. In a corporate network they phish one person and take a workstation. Grey box is not a concession to convenience. It is the accurate threat model for most companies.

When it is the wrong choice. When the question is whether your perimeter is discoverable. Credentials do not help answer that.

White box

Source code, configuration and design documentation are handed over, and testing runs alongside review of the code paths behind the behaviour observed. It finds the deepest issues per dollar, particularly logic errors, cryptographic mistakes, and flaws in code paths that are hard to trigger from the outside.

It costs more because reading code is slow and the tester has to read your language fluently. The output differs in kind. Instead of a list of exploitable findings you get an architectural one, such as authorization enforced per handler rather than centrally, which is worth more than any individual bug.

When it is the wrong choice. When you need evidence for a customer or an auditor that an outsider tested your defence. Some procurement teams read a white box review as the developer marking their own work, even when the reviewer is independent. It is also poor value on a system that is mostly configuration of third-party components, where there is little of your own code to read.

Red team

A red team engagement is not a bigger penetration test. It has an objective, such as reaching a specific dataset or obtaining domain administrator, and the team pursues it across whatever routes are permitted: external systems, phishing your staff, sometimes physical access. Testing runs for weeks, and your detection team is deliberately not warned.

It measures your security operations. Did anything alert. Did anyone notice. How long did it take, and what happened next. A red team report that lists vulnerabilities has missed the point. It should read as a timeline: what was done, when, what fired, what did not, and what the responders did with the signals they got.

When it is the wrong choice. Most of the time. It is the most commonly mis-bought product in security. Without monitoring and someone whose job it is to respond, a red team succeeds on day two, tells you what you already knew, and consumes $50,000 CAD or more that would have bought two thorough application tests. Red teaming is how a mature program finds its blind spots. It is not how an immature one gets started.

The order that works

Application and network testing first, until findings stop being structural. Then detection engineering, so something is watching. Then purple teaming, where the tester and your defenders work in the open to tune alerts. Red team last, when you believe you would catch someone and want to find out if that belief is true. Companies that buy in the opposite order pay the most for the least.

Types described by target rather than knowledge

The other vocabulary in proposals describes what is being tested rather than how much the tester knows. These combine freely with black, grey and white box.

Engagement types by target
TypeWhat it answers
External networkWhat can be reached and exploited from the internet
Internal networkWhat an attacker who is already inside can reach, usually assumed breach
Web applicationWhether your application's own logic can be abused
APIWhether authorization holds when the browser interface is bypassed
MobileClient-side storage, certificate handling, and the backend behind the app
CloudIdentity and access configuration, exposed storage, privilege paths between services
WirelessWhether guest and corporate networks are genuinely separated
Social engineeringWhether people can be induced to hand over access
PhysicalWhether someone can walk in and reach a network port or a desk

Segmentation testing is worth naming separately because PCI DSS requires it explicitly. It verifies that the controls separating your cardholder environment from everything else hold, and it is the one test where a negative result is the whole deliverable.

Choosing without overthinking it

Work down this until a step describes you, then stop. The first match is the answer. The steps below it are what you buy later.

  1. Does a card environment need testing? Buy what PCI DSS prescribes: internal, external and segmentation testing, annually and after significant change. It is the only framework that writes the requirement down, so there is nothing to decide.
  2. Did a customer or an auditor ask? Buy grey box testing of the systems that hold their data, with a retest included. That satisfies SOC 2, ISO 27001 Annex A technical testing expectations and most contract clauses, and it puts the budget where the findings are.
  3. Is the worry a specific component that would end the company if it failed? Add a white box element on that component only, and keep the rest grey box.
  4. Do you have staffed detection, endpoint coverage and closed findings from previous tests? Only then does a red team assessment measure anything, and it costs $50,000 CAD and up.
  5. None of the above, and nobody has asked? Buy an external network test at $6,000 to $15,000 CAD, fix what it finds, and revisit in a year. It is the cheapest honest starting point.

Whatever you land on, get the retest window in writing at the same time. It is the term that decides whether the fixes get verified for free. The detail is on retest and remediation verification.

If none of that resolves it, the test finder asks five questions and routes you to the engagement that fits, and vulnerability assessment versus penetration test covers the other axis: whether a person is involved at all.

If you are testing because you are worried rather than because someone asked, say what you are worried about out loud and scope to that. The best engagement most companies could buy is a grey box application test with a white box element on the two or three components that would end the company if they failed.

The rest of the vocabulary in a proposal, a methodology section or a report is collected on the penetration testing glossary, including the standards a firm should be able to name and the finding types that prove a person tested rather than a tool.

Not sure which type fits

Tell us what you have and what triggered the requirement, and we will recommend a scope before you collect quotes.

Get matched

Common questions

Is black box testing more realistic than grey box?

Only for the first hour of an attack. A real adversary has months and can register an account, phish an employee, or buy credentials from a broker. Your tester has days. Giving them the starting position a real attacker would reach anyway is what makes the comparison realistic, not withholding it.

Will an auditor accept a grey box test?

Yes. Auditors care about independence, scope coverage and evidence of remediation, not about how much information the tester was given. Grey box tests generally produce better coverage evidence than black box ones, because the tester could reach everything in scope rather than only what they discovered.

What is the difference between a red team and a purple team?

A red team works without telling your defenders, and the exercise measures whether they detect it. A purple team runs the same techniques with the defenders watching in real time, tuning detections as they go. If the goal is to improve detection rather than to grade it, purple teaming gets you there faster and costs less.

Do we need internal network testing if everything is in the cloud?

The equivalent is a cloud identity and privilege escalation review rather than a traditional internal test. The question is the same, which is what an attacker who compromises one identity or one workload can reach from there. If you still run an office network with shared file storage and a directory service, that remains in scope regardless of where your product runs.

Can one engagement combine several types?

Yes, and most do. A common shape is external network plus authenticated web application plus a cloud configuration review, quoted as one engagement with a single report. Combining reduces overhead and gives the tester context across surfaces, which is where the better findings come from. What it should not do is reduce the day count for each part.