Kubernetes and container penetration testing
Almost everything sold as container security is image scanning, which is cheap, automatable and worth doing. It is also not a penetration test, and the gap between the two is where every interesting Kubernetes finding lives.
A Kubernetes penetration test in Canada runs about $10,000 to $28,000 CAD for four to eight tester days, and the work that justifies that price is not scanning your images. It is putting a tester inside a pod, as though one of your containers had been compromised through the application it runs, and asking how far they get: whether they can read another namespace's secrets, mount the host filesystem, use a service account token to talk to the API server, or reach the cloud credentials attached to the node. Image scanning answers none of those questions, and a firm quoting $3,500 CAD for container security is selling you the scan.
$10,000 to $28,000 Four to eight tester days, CAD, one cluster
Three products sold under one name
Work out which of these you are being offered. The prices differ by an order of magnitude and so does what you learn.
| Product | What it answers | Effort | Typical cost |
|---|---|---|---|
| Image and registry scanning | Which known CVEs are in the layers you built from | Automated, continuous, runs in the pipeline | $0 to $15,000 CAD a year |
| Cluster configuration review | How far your cluster sits from the CIS Kubernetes Benchmark | Mostly tooling, with a human reading the output | $5,000 to $12,000 CAD |
| Attack-path testing from a pod | What a compromised workload can actually reach | Manual, four to eight tester days | $10,000 to $28,000 CAD |
All are worth money at the right time. The failure is paying for the third and receiving the first, which is the failure described on vulnerability assessment versus penetration test with Kubernetes vocabulary bolted on. Ask the firm how many tester days sit inside the pod, and whether the deliverable includes a reproducible path rather than a benchmark score.
What a tester does once they are inside a pod
The engagement starts from a deliberately weak workload that you and the firm agree on, or from a real application finding if a web application test ran first and found a way in. From there the work follows a predictable ladder, and each rung has a control that stops it.
- Read what the pod was given. Environment variables and mounted volumes hold secrets far more often than anyone intends, and a database password in an env var is readable by anything that can exec into the container or read the process filesystem.
- Use the service account token. Every pod gets one unless you turned that off, and the question is what its role bindings permit. A token that can list secrets cluster-wide is functionally a cluster admin credential.
- Move sideways across the network. Without a network policy, every pod can reach every other pod and every service, including the ones that assumed they were internal and therefore did not need authentication.
- Escape to the node. Privileged containers, hostPath mounts, hostPID, an exposed container runtime socket, or a capability like CAP_SYS_ADMIN each turn a container compromise into a host compromise.
- Take the cloud identity. Once on the node, the instance metadata service and the node's attached cloud role are reachable, and in many clusters that role is far broader than any single workload needed.
Where the findings usually are
Role-based access control is the recurring one. Kubernetes RBAC is permissive by construction, and a real finding is rarely an exotic exploit. It is a service account bound to a ClusterRole because somebody was debugging a deployment on a Friday, a wildcard verb on a wildcard resource, or a role that grants pod creation rights in a namespace that also has a privileged service account, which is a documented route to escalating within the cluster.
Admission control is the second. PodSecurityPolicy was removed in Kubernetes 1.25 and Pod Security Admission replaced it, with the baseline and restricted Pod Security Standards as the profiles you enforce per namespace. Clusters that were built before that change, and never migrated, frequently have no admission control at all, so nothing prevents a deployment from asking for privileged mode and getting it.
What the cloud provider will not let you test
On EKS, AKS and GKE the control plane belongs to the provider. You do not test the API server's own host, the etcd cluster, or the managed scheduler, and attempting to would breach the provider's rules of engagement. What is in scope is everything you configured: your RBAC, your network policies, your admission control, your node pools, your workload identity mapping and your secrets handling. Firms that do this regularly know where that line is, which makes it a useful selection question. The provider-side rules are on cloud penetration testing, and the authorisation wording belongs in your rules of engagement document.
Secrets in Kubernetes are encoded, not encrypted
A Kubernetes Secret is base64 by default. Unless you have enabled encryption at rest for the etcd datastore or you are using an external secrets manager, anything that can read Secrets in a namespace, or read the etcd backup sitting in an object storage bucket, has your credentials in plain text. It is one of the most consistent findings in the category, and a configuration choice rather than a vulnerability, which is why no scanner reports it as one.
Comparing firms for this? Tell us what you need and it goes to the ones in the directory that do this work. No charge, and no phone number required.
Scoping and what changes the price
Cluster count drives more of the price than node count does. Twelve nodes in one cluster is one engagement. Three clusters across development, staging and production, each with its own RBAC and its own cloud role mapping, is three, though two of them go faster than the first. Namespace count matters where namespaces are your tenancy boundary. The test is then whether one tenant's workload can reach another's, the container version of the cross-tenant question that dominates API testing.
| Factor | Added days | Why |
|---|---|---|
| Each additional cluster | 1 to 2 | Separate RBAC, separate identity mapping, separate policy set |
| Multi-tenant namespaces | 1 to 2 | Cross-tenant reachability has to be proven, not assumed |
| Service mesh in use | 1 | Sidecar policy, mesh certificates and mTLS enforcement all need checking |
| Self-managed control plane | 2 to 3 | The API server, etcd and kubelet configuration come back into scope |
| Custom operators and CRDs | 1 to 2 | Operators run with high privilege and are rarely reviewed |
Write all of that into the scope before you collect quotes, using the method on how to write a penetration test scope. Three firms working from the same cluster inventory will quote within a comparable range. Three firms working from the phrase "we use Kubernetes" will not.
When a pentest is premature and a review is the right purchase
The counter-case applies to more Canadian companies than the marketing in this category admits. If you have never applied a network policy, never enforced a Pod Security Standard, run everything in the default namespace, and store secrets as plain Kubernetes Secrets with no encryption at rest, then a penetration test will tell you what a configuration review would have told you for a third of the money. The tester will reach the node on day one, take the cloud role on day two, and spend the remaining days writing up a report whose recommendations you could have read in the CIS Kubernetes Benchmark.
Buy the configuration review, do the work, then buy the test. The test earns its price when there are controls in place for a tester to defeat. That advice costs testing firms money to give, which is why you do not hear it from them.
The other case for not buying: if a customer contract or an auditor is the reason, check what was asked. A clause requiring an annual penetration test of the production environment is satisfied by testing the application and its infrastructure. It does not oblige you to buy a separate Kubernetes engagement. Read the clause, then read SOC 2 penetration testing or the SOC 2 requirements themselves before you scope anything expensive.
Work out which of the three you need
Tell us how your clusters are built and what triggered the request, and we will say whether a review or a test is the right spend this year.
Get matchedCommon questions
Is container scanning the same as a container penetration test?
No. Scanning tells you which published vulnerabilities exist in the software inside your images, which is useful, cheap and should run on every build. A penetration test tells you what an attacker who already controls one of your containers can reach, which depends entirely on your RBAC, your network policy and your node configuration rather than on the contents of the image. A clean scan result and a trivially escapable cluster coexist comfortably.
Can you test a managed cluster on EKS, AKS or GKE?
Yes, with the control plane itself excluded. The provider owns and secures the API server and etcd, and their acceptable use rules prohibit testing them. Everything you configure remains in scope, and in practice that is where the findings are, because managed control planes are not usually the weak part. Expect the firm to raise this before you do.
Should we test staging or production?
Production, if your staging cluster differs in RBAC, network policy or cloud role bindings, which is nearly always the case because staging is where permissions get loosened to make things work. If staging is a genuine mirror built from the same manifests, test it and save yourself the risk conversation. Testing a permissive staging cluster and calling the result a production assessment is the worst of both.
Does a Kubernetes test satisfy our SOC 2 penetration testing evidence?
It can, if the cluster hosts the system in scope for your report. SOC 2 does not name penetration testing as a required control, and auditors accept an independent test of the production environment together with evidence you remediated what it found. What they will not accept is a test of a development cluster presented as a test of production.
How often should we retest a cluster?
Annually as a baseline, and again after any change to how identity works: a move to workload identity federation, a new tenancy model, adding a service mesh, or migrating to a managed control plane. Cluster configuration drifts faster than most systems because deployment manifests change weekly, so pairing an annual test with continuous policy checks in the pipeline is a better use of budget than testing twice a year.