CI/CD pipeline penetration testing
Ask who can deploy to production without a human approving it. In most companies the honest answer is a YAML file that anyone who can open a pull request is able to influence.
Testing a CI/CD pipeline costs roughly $8,000 to $20,000 CAD for three to six tester days, and it is the highest-yield engagement per dollar that most Canadian software companies have never bought. The pipeline is usually the most privileged system in the company. It holds cloud deployment credentials, signing keys, registry credentials and database secrets, it runs code from anyone who can open a pull request, and it is almost never in scope for the annual test because nobody thinks of it as production.
Why the pipeline outranks your servers
A production web server can read the data its application needs. A build runner can deploy any version of any service to any environment. If a tester compromises the first, they have one application. If they compromise the second, they have every application, plus the ability to make the compromise persistent by changing what gets built next time. That asymmetry is the whole argument for scoping the pipeline separately from your application test.
| Compromised system | Reaches | Persists through a redeploy |
|---|---|---|
| One production application server | That application and its data | No |
| The container registry | Everything that pulls from it | Until images are rebuilt from clean source |
| A self-hosted build runner | Every secret injected into any job on that runner | Yes, if the runner is not ephemeral |
| The source repository's workflow files | Every environment the pipeline can deploy to | Yes |
| The cloud role the pipeline assumes | Whatever that role permits, which is usually too much | Yes |
What a tester looks for
A proposal that does not name any of these classes was not written by someone who does this work.
- Workflow triggers that run untrusted code with secrets
- In GitHub Actions, the pull_request_target trigger runs the workflow definition from the base branch but with a token and repository secrets available, while checking out the contributor's code. Combine that with a checkout of the pull request head and you have arbitrary code execution from any fork. This is documented behaviour rather than a bug, which is exactly why it keeps happening.
- Script injection from pull request metadata
- Branch names, commit messages, issue titles and pull request bodies get interpolated straight into shell steps. A branch named to close the quoting and start a new command runs that command on the runner, with whatever the job holds.
- Self-hosted runner poisoning
- A self-hosted runner that is not destroyed after each job keeps whatever the previous job left behind: cached credentials, a modified toolchain, a background process. On a public repository this is a well-known route in. On a private one it is a way for any developer to reach any other team's secrets.
- OIDC trust policies that are too broad
- Federating your pipeline to a cloud role with short-lived tokens is the right design and it replaced long-lived access keys for good reason. The failure is a trust policy that matches on the repository owner but not on the repository, the branch or the environment, so any repository in the organisation, including a new one somebody creates, can assume a production role.
- Secrets that survive in logs and artefacts
- Masking catches a secret printed directly and misses it base64 encoded, split across lines, or written into a build artefact that is retained for ninety days and downloadable by anyone with read access.
- Dependency and build-tool substitution
- Lockfiles that are not enforced, install scripts that run at build time, internal package names that are also claimable on a public registry. The build step is the one place where third-party code runs with your credentials attached.
The controls that answer them
Branch protection with required review, so no single person can merge to a deploying branch. Ephemeral runners, so nothing survives a job. Environment-scoped secrets with a required approval on production. Trust policies that pin the exact repository and reference. Artefact signing and provenance, which is what the SLSA framework's build track is for, with its levels describing progressively stronger guarantees that a built artefact came from the source it claims. Pinning actions and images to a digest rather than a mutable tag.
The compliance angle nobody uses
Your pipeline is where your change management evidence comes from. Auditors ask how a change gets from a developer's laptop to production, who approved it and whether the approval can be bypassed. If a tester can show that branch protection is enforced on the interface but not on the API, or that an administrator can push directly, that is not only a security finding, it invalidates a control you have been describing to your auditor. There is more on what auditors accept at SOC 2 requirements and on SOC 2 penetration testing.
Comparing firms for this? Tell us what you need and it goes to the ones in the directory that do this work. No charge, and no phone number required.
This is a review, and firms should say so
Most of a pipeline engagement is reading configuration: workflow definitions, trust policies, runner setup, branch protection rules, registry permissions. There is real exploitation in it, usually a proof-of-concept pull request from a throwaway account that demonstrates code execution on a runner, but the ratio is perhaps one day of attack to four days of review.
That is the correct shape for the problem. The weaknesses are configuration choices rather than software defects, and no black box probing finds a trust policy you cannot see. White box access is mandatory. A firm offering to test your pipeline without read access to the workflow files and the cloud trust policies is offering something much less useful, and the reasoning is on penetration testing types.
Check your own pipeline before you buy anything
Most of these you can answer in an afternoon. If enough come back badly, fix those first and buy the engagement afterwards, when a tester will spend your money on what you could not find yourself.
0 of 0 confirmed ·
Scoping and price
| Activity | Days | Amount |
|---|---|---|
| Workflow and repository configuration review | 1.5 | $3,000 |
| Cloud trust policy and permission analysis | 1.0 | $2,000 |
| Runner and registry testing, proof of execution | 1.5 | $3,000 |
| Secrets handling and artefact review | 0.5 | $1,000 |
| Reporting and remediation ordering | 1.0 | $2,000 |
| Typical mid-size engagement | 5.5 | $11,000 |
What moves the number is repository count and the number of distinct deployment targets, not lines of code. Ten repositories that all use one shared reusable workflow is close to one engagement. Ten repositories that each invented their own deployment is close to ten. Write that down before collecting quotes, using how to write a penetration test scope, and check the proposals against what a quote should itemise.
When not to buy this
If you deploy by hand, from a laptop, with no pipeline, there is nothing here to test. Your problem is the absence of change control, which is a governance question and a cheaper one to solve. Fractional security leadership handles that better than a testing firm does, and vCISO services is the shape of that purchase.
If you have a pipeline but have never turned on branch protection, skip straight to the checklist above and spend the money on fixing what it surfaces. Paying $12,000 CAD for someone to tell you that anyone can push to the branch that deploys production is not a good use of the budget. Buy the test when the controls exist and the question is whether they hold.
Get the pipeline into your next scope
Tell us how you build and deploy, and we will say whether it belongs in the annual test or needs its own engagement.
Get matchedCommon questions
Is CI/CD testing included in a normal penetration test?
Almost never, unless you asked for it in writing. A standard application or network test scopes the running system, and the pipeline that built it sits outside that boundary. If you want it covered, name the repositories and the deployment targets in the scope document, because a firm will not add a system it was not told about.
Do we need to give the testers access to our source code?
Yes, read access to the repositories and to the workflow configuration, and ideally read access to the cloud trust policies. The weaknesses in this category are configuration you cannot see from outside, so a black box engagement would spend most of its budget guessing. Scope it as a white box review and accept the access requirement as part of the purchase.
What if we use GitLab, Jenkins or Azure DevOps rather than GitHub?
The categories transfer directly even though the names change. Every platform has a trigger that runs contributor code, a way to attach secrets to a job, a runner or agent model, and a mechanism for federating to a cloud role. Self-hosted Jenkins usually tests worse than the managed platforms because plugins accumulate and the controller often runs jobs itself.
How does this relate to a software supply chain assessment?
Pipeline testing is the part of supply chain security you control directly. A wider supply chain assessment also covers the dependencies you consume and the vendors you buy from, which is a procurement and questionnaire exercise rather than a testing one. Companies usually get more from securing their own build first, because that is where an attacker gets use over everything downstream.
How often should a pipeline be tested?
Annually alongside your other testing, and again whenever you change how deployment authenticates: moving to OIDC federation, adding self-hosted runners, introducing a new deployment target, or consolidating onto shared reusable workflows. Those changes are exactly where the trust policy mistakes appear. How often you should pentest covers setting a cadence across all your systems.