What Is a Penetration Test, and When Does a Company Need One
A scan, an audit and a penetration test are three different jobs at three different prices. Which of them answers your question, and what UK law actually requires.
A scan, an audit and a penetration test are three different jobs at three different prices. Which of them answers your question, and what UK law actually requires.
You ask for a quote for website security and receive three that differ tenfold, and one says “security scan”, the second “security audit”, the third “penetration test”, but all three carry the same word “security”, which says nothing about what actually differs. They are not one job in three grades and not three prices for the same work, but three different jobs that answer three different questions, so the first step is not comparing prices, but understanding which of those three questions you actually want answered.
This article answers two: what a penetration test is, and whether it is needed for your company specifically. The second question is the more important one, because for most small and medium-sized companies no statute names a penetration test as a duty, and a quote that stays silent about that is selling you the right service at the wrong moment.
We write this while selling a website security audit, and there is no contradiction in that, because an audit that has room for a manual check, and a penetration test that a regulation requires, are not the same job, and a company that buys the second when it needed the first pays more and learns less.
Three jobs sold under similar names
Industry terminology is messy here, but under it there is a clear boundary that a standards body names best: NIST SP 800-115 defines a penetration test as security testing in which assessors mimic real attacks in order to find ways around a system’s defences, and notes that it looks for combinations of vulnerabilities, not isolated findings. The same publication describes a vulnerability scan as a technique that identifies hosts and the known vulnerabilities that correspond to them.
In practice that means four distinct jobs, each worth naming separately. A vulnerability scan is automated, and in it a tool compares your system with a database of known vulnerabilities and returns a list. A vulnerability assessment is that same list, which a person has checked and ranked, typically without exploiting any of the vulnerabilities found. A penetration test is human-led work in which the weaknesses found are exploited and chained, in order to establish how far an attacker really gets. A compliance audit answers an entirely different question: whether the system meets a named standard or legal norm.
The boundary between the first three is drawn most clearly by the payment-card standard, which does it not with a definition but with a purpose: the PCI Security Standards Council’s guidelines distinguish a penetration test from a scan by aim: the scan identifies, ranks and reports vulnerabilities; the test looks for ways to exploit them in order to bypass the system’s defences. The same council adds both edges in its standard: a scan by itself is not a penetration test, and a test that only tries to exploit the scanner’s findings is not sufficient either.
So the question “do you need a penetration test” is not a question about budget, but about what you want answered: whether your system has known weaknesses, or whether someone can actually do something with them.
What each of them answers, and what it does not say
A scan answers quickly and cheaply, and its weakness is context, because the tool does not know which of a hundred flagged rows is on your payment form and which is in a test environment that nobody can reach from the internet, and it also does not know that two separately harmless faults together give access to the database — so the list is the start of the work, not the result of the work.
A penetration test answers more slowly and more expensively, and its value is precisely in the chain, because the tester is looking not for “is there a vulnerability here” but for “what can I do with it” — whether from a public form one can reach administrator rights, whether from one customer’s account one can see another customer’s data, whether from a test environment one can reach the production database. The answer to that question is the one a director understands without a translation.
A compliance audit does neither, because it checks whether the system meets a norm, and its result is a verdict on conformity, not a list of technical weaknesses; that is exactly why a company can pass a compliance audit and be breached the same week, and there is no contradiction in that, because two documents answer two different questions.
None of these jobs replaces the others and none of them is “better” than the others, so the only meaningful question is which of them currently answers what you actually need to know.
The confusion of names is not only a local problem, and you see that by comparing markets: in Finland, Sweden and Norway the market sells two jobs — scanning and a penetration test — and the middle option has no name; it is treated as the result of a scan, not as a separate service. In Germany it is the opposite, because there one word often covers all three jobs, including the test itself, and that is precisely why, when quotes are compared, the name is the weakest possible guide, and the only safe question remains whether the vulnerabilities found will be exploited.
Whether the law requires it of you specifically
Precision is worth it here, because this is the place where quotes most often overclaim: in the United Kingdom no statute names a penetration test as a duty on every company that has a website.
The Network and Information Systems Regulations 2018, which came into force on 10 May 2018, set duties for operators of essential services: appropriate and proportionate measures to manage risks to the network and information systems on which the essential service relies, and to prevent and minimise the impact of incidents. A penetration test is not named in any regulation of the instrument itself.
What names testing is the NCSC Cyber Assessment Framework, principle B4, which competent authorities use to assess those operators. It lists “regular vulnerability and security assessments, e.g. penetration tests and vulnerability scans” among the well-known methods, but only for the network and information systems that support the essential function. The same regulations also let a competent authority serve an information notice requiring the information it reasonably needs in order to assess the security of those systems.
Those are two conditions together, not one: a company that is not an operator of essential services, and not a relevant digital service provider, picks up no duty from this point, and an operator whose particular system does not support the essential service also does not, so the first step is not asking for a quote, but a self-assessment of whether you are in scope at all.
How to tell whether you are in scope
The NIS Regulations do not list companies by name, but describe sectors and size, and the duty to work out your own status sits on the company itself, which in practice means two questions: whether your activity falls within one of the kinds of essential service set out in the regulations, and whether you meet the threshold described for that kind. Both have to be answered by you, and the answer has to be documented — which is why the first expense in this area is usually legal advice, not a technical service.
If the answer is “no”, the further points about which systems support an essential function do not apply to you at all, and if the answer is “yes”, the second question is which of your systems the essential service actually relies on. The expectation of a penetration test is attached to those systems, not to in-scope status as such, and a company can be an operator of essential services none of whose systems this expectation applies to.
There is also a third path, which still requires you to be in scope but does not depend on a particular system supporting the essential function: a competent authority may serve an information notice on an operator of essential services requiring the information it reasonably needs in order to assess the security of that operator’s network and information systems, in a particular case. That is not a duty you can plan for, but it is a reason to know who on your side would answer such a request.
The other thresholds name testing, not a penetration test
Article 32 of the UK GDPR requires a process for regularly testing, assessing and evaluating the effectiveness of security measures, and the scale of the requirement is tied to risk. That is not the same as a duty to commission a penetration test once a year, and a quote that presents Article 32 as such restates the regulation more loosely than it is written. The Information Commissioner’s Office names vulnerability scanning and penetration testing as techniques with which you can do that testing; the duty to show that testing happens at all is real; the duty to choose this particular kind of testing does not follow from Article 32.
Requirement 11.4 of the payment-card standard PCI DSS names a penetration test directly, every twelve months and after significant changes, but that is a contractual duty and it applies to how you accept cards. An online shop that routes payments entirely to a payment service provider and does not itself process card data, and everyone else whose environment card data enters, answer this requirement differently. That is a question your acquirer answers, not an article.
In the financial sector, CBEST is a threat intelligence-led penetration testing framework run by the Bank of England, the Prudential Regulation Authority and the Financial Conduct Authority, and they reserve it for the firms they select. The standard ISO/IEC 27001 requires vulnerability management and security testing, without naming a penetration test; the claim that a certificate is refused without one is widespread and is not in the standard.
If none of these thresholds applies to you, then you have no statutory duty to commission a penetration test, and that does not mean there is nothing to do, but that the work to be done is different.
If no threshold has been crossed
For a company that is not in scope, does not process card data and does not operate in the financial sector, the meaningful work is usually what happens regularly, not once every few years. That is updates that have an owner and a deadline; backups that someone has restored at least once and verified that they really restore; multi-factor authentication (MFA) for administrator accounts, which Cyber Essentials requires of those who certify to it and which is equally useful for everyone else; and regular scanning whose results someone actually reads.
This is not a lesser form of an answer, but a different kind of work. Most of the compromises we work with do not start with a sophisticated attack, but with an unpatched component or a password that also worked somewhere else, and a penetration test that happens once every few years does not protect against that. If your website has already been hit, the sequence is different and a separate article describes how to recover a hacked website.
A penetration test becomes justified when there is something that can be lost, and when the scale of the loss is larger than the price of the test: a system that holds other people’s data, an integration that touches money, or a client who demands proof. Until that moment it is the right job in the wrong order.
What happens during the test
The work starts with scoping and ends with a report, and between them there are three stages worth understanding before quotes are compared, because they are exactly what explain why one test costs what it costs, and why another for the same system costs ten times less.
The first stage is reconnaissance, in which the tester collects everything that can be learned about the system from the outside — which addresses are public, which technologies and versions are visible, where the login forms are, which files are available without authorisation. Nothing is exploited at this stage yet, but this is where findings nobody expected most often sit: a forgotten test environment, an open directory listing, a backup lying at a predictable address.
The second stage is the test itself, in which automated tools are run so that the known is not missed, but the decisions are made by a person who checks whether a finding is real, tries to exploit it and looks at what is gained. This is where the chain appears — access to one account, from that to a function that should not have been reachable, from that to data. Separately none of the steps is dramatic; together they are a story.
The third stage is proving and writing it down, because every finding has to leave evidence that can be repeated: which request was sent, what the response was, what changed. Without that the report is an opinion, and the developer who receives it spends a day trying to understand what the tester actually saw.
That is also why the timescales are what they are: a test of one small website is a few days, but a system with several roles, integrations and payments is weeks, and a quote that promises a full penetration test in a single day is describing not a test but a scan.
Why quotes differ tenfold
When two quotes for the same website differ by a factor of ten, the difference is almost never in the margin, and it is almost always in scope and method: one offers an automated scan with a tool that is launched in an hour and whose report the same tool generates, while the other offers a week of human work in which the tools are only the start, and it is exactly this difference that the name “security check” conceals in both cases.
The second price-forming factor is the complexity of the system, and that can be judged before the conversation starts: one public website without user accounts is one job, but a system with several roles, payments, an external integration and data that belongs to customers is an entirely different one, because each role is a separate boundary to be tested and each integration is a place where two systems trust each other more than they should.
The third is what you receive after the test and how long the provider stays: a report without priorities, without evidence and without a re-test costs less because it is less work, and for a company that then has to fix the findings, exactly these three things decide whether the document becomes a task list or a folder nobody opens again.
So comparing quotes by price is possible only when the scope has been written the same way, and the simplest way to achieve that is to write the scope yourselves and ask everyone to quote against it, rather than letting each provider define their own.
When a test goes stale
The result of a penetration test describes a particular system on a particular date, which is obvious, and yet this is exactly where most of the misunderstanding between provider and client arises, because a report that is a year old describes code that has since changed dozens of times.
The written intervals admit this themselves: the CAF speaks of “regular” assessments rather than a named calendar, and in the payment-card standard next to the yearly interval stands “after significant changes”, so the calendar, where it stands, is only a minimum, and the real reason to test is a change.
In practice that means a new reason to test is a change that changes the attack surface: a new public function, a new integration with an external system, a change of authentication, a move to different hosting, a new user role with wider rights. A change of colours or a text correction does not become such a reason, however visible it is.
The other reason is that the surroundings have changed, not your code: a vulnerability in a framework you use is disclosed after the test, and a test that did not mention it was not wrong, because at that moment it did not yet exist. That is why regular scanning and an update process are what happen between tests, and the test does not replace them.
Without the owner’s authorisation the same acts are unlawful
A penetration test does not differ technically from an attack, and the only thing that distinguishes them is a document: the owner’s authorisation, in which the scope, the time and the boundaries are named, and which is worth putting in writing even when the statute does not require that form. Without it the same acts are the same acts, and that has legal consequences — in the United Kingdom, unauthorised access to computer material is an offence under section 1 of the Computer Misuse Act 1990.
In practice that means the authorisation is given by the person the system belongs to, not the person who maintains it. If your website runs on a hosting service, the provider also has to know that the test will happen, because otherwise their defences will stop it or lock your account. If the system contains a third-party component you do not control, it is not in scope.
The boundaries of the scope have to be written before, not after, and that list includes which addresses are in scope and which are not, whether we test the production environment or a copy, what happens if the test interrupts the service, and who on your side is reachable at night; this conversation takes an hour and resolves most of the disputes that would otherwise arise in the middle of the test.
What a test is not: red teaming, a bug bounty and a compliance check
Beside a penetration test there exist several jobs that tend to be called the same thing, and the differences between them are not academic but practical — they decide what you commission and what you receive.
A red-team exercise checks not the system but the defence: whether your people and processes notice an attack and what they do. The scope is wider, the duration longer, and part of the value is precisely that the defending side does not know that an exercise is taking place. For a company that has nothing to notice, because nobody reads the logs, this work is premature.
Bug bounty is a model, not a test: you publish rules and pay for findings to those who send them. It can find what one tester missed, but it gives neither a guarantee of coverage, nor a deadline, nor a report that can be attached to tender documents.
Threat-led penetration testing (CBEST) is a separate, supervised piece of work in the financial sector, and it is defined by the Bank of England, the PRA and the FCA. If your company is not a firm they have selected, this term does not belong in your quote.
Another boundary that tends to be erased is the one between “black box”, “white box” and “grey box” — how much the tester already knows about the system at the start. That is an industry convention, not a statutory requirement, and it has a direct effect on price and on what the test will find. A tester without access mimics a stranger; a tester with an account and documentation reaches further in the same time. Neither variant is the more correct one; the question is what you are afraid of.
What you receive and how to read it
The result of a test is a report, and its value is in the priorities, not in the number of findings, because a report with a hundred rows that does not say which to start with is exactly as unusable as a scanner printout. A good report tells, for each finding, what an attacker can do with it, how easily, and what specifically to change.
The second thing to ask for is a check after the fixes, because a finding that has been fixed, and a finding someone thinks has been fixed, differ from each other, and the only way to establish that is to test again. We do that within thirty days of the fixes, and this deadline is worth asking of every provider.
The third is what your developer will do with the report. A finding described with a CVE number and without context means the developer has to go looking; a finding that has the particular request and the place in the code attached means a fix. If one company does the development and another the test, this difference is the one that decides whether the fixes will be made in a week or in a quarter.
The fourth is what must not be in the report: a claim that the system is now secure. A test shows what, in the agreed scope on the agreed date, it was possible to do. It does not prove that there is nothing else, and a provider who promises that is selling you comfort.
There is one more reason why this conversation is worth starting earlier than it seems necessary: a test that happens a week before the system goes live finds the same things it would have found three months earlier, but there is no longer time or budget to fix the findings, and in practice it ends with a list that is accepted as a risk rather than with fixes, so testing before going live is worth doing for that reason, including for those whom no statute binds.
And the last thing worth saying directly: a penetration test is not a certificate that the system is secure, but a certificate that on a particular date in a particular scope a known skill did not find a way past something particular, and that is why the most valuable part of the report is often not the list of findings, but the description of what was tried and failed, because that is the part that tells the next tester in three years where it is not worth starting from zero.
What to prepare before the conversation
For a quote to be comparable at all, the provider has to know the scope, so prepare a list of the addresses and systems that belong in it, say whether we are testing the production environment or a copy, say what in the system must not be touched, and name a person who can authorise stopping the test.
Production or a copy
This is the question that decides both price and risk, because a test in production shows what is actually reachable, and that is exactly why it can break something: overload, fill the database with test records, send real emails to customers, or stick in a defensive system that blocks the tester and then also blocks some of your users.
A test on a copy is safer and at the same time less complete, because a copy is rarely identical: it often lacks the real integrations, the real volume of data and the real configuration, and it is often in the configuration that the problem sits. If you choose a copy, write down how it differs from production, because that list is also a list of what the test did not check.
The middle path we use most often: read operations in production, writing and potentially destructive ones on a copy, with a window agreed in advance and a person who can stop it. That is not a compromise for the sake of price, but a way of getting the answers of both variants without interrupting the work.
The opposite list is also useful, namely what is not in scope, because silently assumed boundaries are the ones later fought over. Third-party services you do not maintain do not belong in it, and testing them without those parties’ authorisation is the same problem the previous section is about. If your website uses an external payment window, an external chat window or external analytics, those are other people’s property, and a quote that offers to “check those too” is offering what it must not.
Finally say what will happen with the findings afterwards: who will fix them, in what time, and whether the provider will check again after the fixes. A test without this agreement often ends with a document nobody opens, and that is the most expensive possible version: paid for knowledge that is not used.
Also say what answer you are looking for, because “we have to meet a requirement” and “we want to know whether someone can get to customer data” are two different jobs at two different prices, and a provider who does not ask which of them is yours will offer the one that is more convenient for them.
If you are not sure which side of the threshold you are on, that is worth starting with. The automated audit usually answers the question about known weaknesses more cheaply and more quickly than manual expert work, and its result also says whether a penetration test is the next step. A conversation about scope is worth starting with a description of the process, not with a list of technologies, because the scope is set by what you lose if the system fails.
Frequently asked questions.
What is a penetration test?
A penetration test is a human-led security check in which the tester, with permission, mimics a real attack, exploits the vulnerabilities found and chains them, in order to establish how far an attacker can really get. NIST SP 800-115 defines it as testing that looks for ways around a system’s defences and for combinations of vulnerabilities, not isolated findings. You tell it from a scan by the result: a scan returns a list, a test returns an answer to the question of what can be done with that list.
Is a penetration test mandatory?
For most companies, no. In the United Kingdom no statute names a penetration test as a duty on every company with a website. The NIS Regulations require appropriate measures of operators of essential services and relevant digital service providers, and do not name a penetration test in the regulations themselves; the NCSC CAF, which competent authorities use to assess those operators, names penetration tests and vulnerability scans as examples of regular assessments, and only for the systems that support the essential function. Article 32 of the UK GDPR requires regular testing commensurate with risk, without naming a penetration test, and PCI DSS requirement 11.4 applies to how a company accepts payment cards.
How does a penetration test differ from a vulnerability scan?
By whether the weaknesses found are exploited. A scan is automated and compares the system with a database of known vulnerabilities. A penetration test is an active process in which the vulnerabilities found are typically exploited — that is how the PCI Security Standards Council puts it, and it also adds the other edge: a scan by itself is not a test, and a test that only checks the scanner’s findings is not sufficient.
How often should a penetration test be done?
If the expectation comes from the NIS Regulations and the CAF, then regularly, on a cadence set by risk and by change, not by a named calendar in the regulations. If the duty comes from PCI DSS, then every twelve months and additionally after significant changes to the infrastructure or the application. If there is no statutory duty, frequency is set by the pace of change: a test carried out before two large rebuilds describes a system that is no longer there.
May a penetration test be carried out without the system owner’s permission?
No. A penetration test does not differ technically from an attack, and the only thing that distinguishes them is the owner’s written authorisation with a named scope, time and boundaries. Unauthorised access is an offence under the Computer Misuse Act 1990. The authorisation is given by the person the system belongs to, not the person who maintains it, and the hosting provider also has to know about the test, otherwise their defences will stop it or lock the account.
Security audit. We find the holes before hackers do — OWASP Top 10, a manual penetration test, a report with priorities.