GetPentest

Penetration testing methodology

A methodology is the order the work happens in and the reason each step exists. Knowing it is how you read a proposal and tell structured testing from a scan with a cover page.

Last reviewed 2026-09-16Written by Jacob Masse, TrazTech Inc.

Every proposal you receive will name a methodology, usually several, in a row near the bottom of page two. The documents behind those names are real, but a citation is not evidence that the work followed one. The way to tell is to know what each phase produces, what the tester needs from you, and what the deliverable looks like when it is done. This page sets that out phase by phase, then says what the named methodologies contain.

The phases of an engagement

A professional test runs in the same order whether the target is a web application, an internal network or a mobile app. The proportions change, the sequence does not, because each phase depends on the output of the last.

Engagement phases and what the client supplies
PhaseWeight in a typical engagementWhat the client provides
Scoping and rules of engagementBefore the clock startsAsset list, exclusions, window, contacts, signed authorisation
Reconnaissance and enumerationModest, but it sets up everything after itArchitecture notes, domain and IP ownership confirmation
Vulnerability identificationSubstantialScanner allowlisting, a stable test environment
ExploitationSubstantialAn escalation contact who answers within the hour
Post-exploitation and lateral movementVaries widely, and only where scope allows itWritten agreement on how far the tester may go
Business logic and authorisation testingVaries by scopeTwo or more accounts per role, and a product walkthrough
Evidence capture and reportingVaries by scopeFactual review of the draft
RetestA separate block, weeks laterFixes deployed, and a list of what changed

Scoping and rules of engagement

Scoping happens before anyone touches a keyboard, and it is where the engagement is won or lost. The output is a written document, agreed by both sides, answering five questions.

  • Targets. Exact hostnames, IP ranges, application URLs, API base paths, mobile builds, cloud account identifiers. "The platform" is not a target. A list is.
  • Exclusions. What must not be touched, and why: shared hosting a third party owns, a payment processor's endpoints, a legacy host running the one process nobody can restart. Denial of service testing is excluded by default in almost every commercial engagement.
  • Testing window. Dates, hours and time zone. Business hours finds real behaviour under real load. Overnight protects a fragile system at the cost of realism.
  • Escalation contacts. Two names with phone numbers on both sides, for the case where the tester finds evidence of an existing compromise, causes an outage, or reaches something nobody expected to be reachable.
  • Authorisation. A signed statement from someone with the authority to grant it, naming the targets and the window. Where the systems sit in a cloud or hosting tenancy the provider's testing policy applies on top, and the client confirms the targets are theirs to authorise.

Scoping also sets the price, because the price is a count of days and the days come from the asset list. What belongs in it is on penetration testing rules of engagement, and building the asset list is on how to write a penetration test scope. Starting from nothing, the scoping questionnaire produces that list in the shape a tester expects.

Comparing firms for this? Tell us what you need and it goes to the ones in the directory that do this work. No charge, and no phone number required.

Reconnaissance and enumeration

The tester maps what exists before deciding what to attack. Passive reconnaissance uses sources outside your systems: certificate transparency logs, public DNS records, code repositories, job postings naming your stack, breach corpora holding credentials for your domain. Active enumeration then touches the in-scope targets for live hosts, open ports, service versions, subdomains and API endpoints.

The output is an attack surface inventory, often the first genuinely useful deliverable. Clients routinely learn here that a staging environment is public, or that an API version believed retired still answers. What you provide is confirmation of ownership, because the tester will find hosts resembling yours and cannot touch them without your word that they are in scope.

Vulnerability identification, automated and manual

This phase produces candidate findings. Automated discovery runs first because it is cheap: a scanner compares what it sees against a signature database and reports missing patches, outdated libraries, weak TLS configuration, exposed administrative interfaces and known CVEs. It is fast, wide, and better than a person at what it does.

Then a human verifies every candidate, which is the step that separates the two products. Scanners infer from a version banner, and a banner lies whenever a distribution backports a fix without changing the version string. The tester reproduces the condition, confirms it is real in your configuration, discards the ones that are not, and adjusts severity to your context. A critical finding on a host reachable only from a management network is not critical.

Manual identification also covers what no scanner has a signature for: injection points reached only through a multi-step workflow, file upload handling, and session and token logic. A quote that describes coverage in tool names rather than tester days has priced the first half of this phase and not the second, which is the distinction drawn on vulnerability assessment vs penetration test.

Exploitation, and what responsible means here

A weakness that has not been exploited is a theory, and theories get deprioritised in sprint planning. Responsible exploitation means proving impact without causing harm, which in practice is a set of rules a professional applies unprompted.

  • Prove read access rather than write access. Retrieving one record you should not see proves the flaw. Modifying it proves nothing more.
  • Extract the minimum. Where a database is reachable, the evidence is a version string, a table list and one masked row.
  • Stop at the scope boundary, and record where, so the report can say what was left untested.
  • Run no exploit known to crash the service against production without written agreement in the rules of engagement.
  • Clean up, and list web shells, added accounts, uploaded files and scheduled tasks with timestamps so your team can confirm removal.
  • Call the escalation contact on evidence somebody else got there first. That is an incident, not a finding, and the engagement pauses.

Post-exploitation, privilege escalation, lateral movement

Access on its own is a weak finding. What turns it into business risk is what it reaches: what data is readable, what credentials are cached on the host, what other systems trust it. Privilege escalation is the move to higher privileges on the same system through a misconfigured service, a writable path in a privileged process, an over-permissive cloud role or a token with more scope than the feature needed. Lateral movement is the move across systems, using recovered credentials, trust between accounts, or a segment that turned out not to be segmented.

Both are in scope only when the rules of engagement say so, and the limits are written before testing starts: whether domain administrator is a goal or a stopping point, whether production data may be touched, whether persistence may be established. An internal engagement without this phase is an internal vulnerability scan. The engagement shapes are set out on penetration testing types.

Business logic and authorisation testing

This is the phase that earns the fee, and the one quietly dropped when a quote is too low to contain it. No tool can do it, because no tool knows what your application was built to allow.

Authorisation testing asks whether the rules enforced in your interface are also enforced in your API. The tester holds two accounts in the same role at different customers, and two at different privilege levels in one customer, then takes an identifier visible to one and requests it as the other. Across every object type and verb that is slow, repetitive, and exactly where multi-tenant products fail.

Business logic testing asks whether a sequence of individually legitimate actions produces a result you never intended: skipping an approval step by calling the final endpoint directly, applying a discount twice, cancelling a transaction after the fulfilment trigger fires, or registering with an address that reuses a pending invitation and inherits its permissions.

These findings are described in your own vocabulary, naming your roles, objects and workflow. A report containing none of them is a report about software in general rather than about your product, which is the first test to apply when reading one. More is on how to read a penetration test report.

What the client provides here matters more than anywhere else: working accounts in every role, two tenants where the product is multi-tenant, and forty five minutes of a product person walking the tester through what the system is supposed to do. Testers given that find more, and it is the cheapest improvement available to a buyer.

Evidence capture

Evidence is captured continuously, not written up at the end from memory. For each finding the tester records the request and response pair, a screenshot, the timestamp, the account used and the reproduction steps, with sensitive values masked in the report. That way engineers reproduce the issue without a meeting, an auditor sees the test happened rather than taking a summary on trust, and the retest has something precise to re-run. Where the evidence is held and under whose jurisdiction is a fair scoping question, covered on penetration testing data residency in Canada.

Reporting

The report has two audiences and a serious one serves both. The executive summary states what was tested, what was reachable, what the business consequence would be and what to do first, in language an executive can act on without a translator. The technical body carries each finding with justified severity, affected components, reproduction steps, evidence, and remediation guidance naming your technology rather than a generic control. Where a scoring system is cited, the version and vector belong in the finding.

It should also state what was not tested and why, the section that tells an experienced reader the work was real. A draft goes to you for factual review before the final issues, because a correction at draft stage is cheaper than a wrong finding circulating with customers.

Retest

The engagement is finished when the findings are closed and somebody independent has confirmed it. Retest re-runs the reproduction steps for each remediated finding and records the result: fixed, partially fixed, not fixed. Partial fixes are common, because a fix at one endpoint misses the others that share the flaw.

Ask during scoping whether retest is included, how long the window stays open, and whether the result is a fresh report or a letter. Many buyers find out after remediation that it is a separate purchase. What to expect is on retest and remediation verification.

The methodologies named in proposals

These documents are public and free to read. They cover different things, and a proposal listing all of them without saying which parts apply to your scope has told you nothing.

What each named methodology actually covers
NameWhat it isWhat it is good for
OWASP Web Security Testing Guide A testing manual for web applications, organised by category, with individual test cases The strongest reference for web test coverage. Ask which sections were applied
OWASP Top 10 An awareness list of the ten most critical categories of web risk Communication and training. Not a methodology, and a test scoped only to it is thin
OWASP API Security Top 10 The same idea for APIs, with broken object level authorisation at the top Framing an API engagement, because it names the failures that dominate API findings
PTES The Penetration Testing Execution Standard, defining engagement phases from pre-engagement through reporting Process and expectations. It says how an engagement runs, not which tests to perform
NIST SP 800-115 A US federal guide to security testing, covering planning, execution and post-execution Program structure and a vocabulary auditors recognise. A framework rather than a test catalogue
OSSTMM The Open Source Security Testing Methodology Manual, a formal model for measuring operational security across channels, human and physical included Measurement and repeatability, and testing beyond the technical channels

Read the pairings. A web engagement should cite the Web Security Testing Guide for coverage and PTES or NIST SP 800-115 for process. An API engagement should cite the API Security Top 10 and say how authorisation was tested per object type. One citing only the OWASP Top 10 has named an awareness document as its test plan. The follow-up question is simple: which sections applied to our scope, and which did you exclude.

Black box, grey box and white box

These describe how much the tester knows at the start.

  • Black box. No credentials, documentation or source. A realistic picture of the external perimeter, and paid days spent rediscovering what you could have written on one page.
  • Grey box. Credentials for each role, architecture notes, an API specification and a product walkthrough, but no source. The tester spends the engagement inside the application, where the authorisation and business logic flaws live.
  • White box. Everything, source and configuration included. The highest coverage per day for one component, and the right choice for a payment flow, a cryptographic implementation or an authentication service.

Grey box is usually the best value. The findings that cause real losses are reached from inside an authenticated session, and a black box tester spends a third of a short engagement getting to where a grey box tester starts on day one. Withholding credentials buys a harder test of the login page, not a better test of the product. If the question is whether your monitoring would catch an intruder, that is a red team objective and a different engagement.

Which shape fits can be worked through with which pentest do I need, and the other buyer tools cover cost and severity triage. What these phases do to a quote is on penetration testing cost in Canada.

Compare firms on method, not on page count

Tell us what triggered the requirement and what the system does, and we will put it to firms that answer in days and phases.

Get matched

Common questions

Is there one standard methodology every firm follows?

No, and no body enforces one. The published documents split between process standards, which say how an engagement runs, and testing guides, which list what to test. Serious firms combine one of each and say which sections applied to your scope. The signal is specificity about your engagement, not the number of standards listed.

How long should each phase take?

There is no standard split, because scope drives it. A product with one user role and forty screens and a product with four tenants and a partner portal need very different amounts of authorisation testing, and that phase is usually the largest single block on a web engagement. What matters is that the proposal states a split at all. Ask any firm how many days it has allowed for each phase, and whether reporting and retest sit inside that number or outside it.

What happens if the tester breaks something?

The escalation path activates: the tester stops, calls the named contact, and documents what was run and when so your team can diagnose quickly. That is why both sides name two contacts before testing starts. Ask which entity signs the contract, under which province's law, and what insurance sits behind it.

GetPentest is operated by TrazTech Inc.