AI has not created a new class of web application flaw. It has made the familiar ones cheaper to find and quicker to exploit at scale. The NCSC set this out in its assessment of the near-term impact of AI on the cyber threat, published on 24 January 2024, judging that AI will almost certainly increase the volume and impact of attacks over the following two years. For web application security testing, the shift is mostly about timing.

What AI changes for attackers
AI lowers the cost of the reconnaissance and scripting that used to slow attackers down. The NCSC assessment is careful about this: the uplift is greatest in reconnaissance and social engineering, weakest in genuinely novel exploitation. In practice that means the gap between a proof of concept appearing on GitHub and a working exploit hitting your login page has shrunk from days to hours. It also means the badly worded phishing email is gone. Source maps left in a production JavaScript bundle used to take an analyst an afternoon to unpick. Feed the same bundle to a model and the internal endpoint list comes back in a minute, parameter names included.
Features built on language models are now attack surface
If your application calls a language model, that model is part of your attack surface. Prompt injection sits at the top of the OWASP Top 10 for Large Language Model Applications, and the indirect variety causes the most trouble: a user uploads a document, the assistant summarises it, and instructions buried in the document are executed with the assistant’s privileges. The same patterns come back again and again. A support chatbot wired to an orders API through a service account that can read every customer. Model output rendered as HTML without encoding, which turns a summary into stored cross-site scripting. Retrieval systems that index documents from one client area and answer questions about them from another. Debug routes that return the full system prompt, including the key somebody pasted into it.
“The chatbot findings we report are rarely about the model itself, they are about what the developers connected it to. If your assistant can call an internal API using a service account that sees every customer record, then a prompt filter is the only thing standing between a stranger and your database. Give the model its own least-privilege credentials and enforce authorisation at the API, never in the prompt.”
William Fieldhouse, Director, Aardwolf Security Ltd

Why your web application security testing window matters more than it did
Shorten the gap between a change going live and the test that covers it. An annual test made sense when your application changed twice a year. If you deploy weekly, an annual snapshot describes an application that no longer exists by the spring. Two adjustments work without inflating the budget. Book the full assessment annually but hold days in reserve for significant releases, and put authenticated scanning into the pipeline so routine issues are caught before a consultant sees them. Manual web application penetration testingthen goes where it earns its money: authorisation, business logic and anything a model touches.
What to ask a testing provider now
Ask whether the provider tests language model features by hand, and how they deal with responses that differ between attempts. It is a fair question and the answer is revealing. Non-determinism breaks the usual evidence model, so a tester has to repeat each attack across fresh sessions and record full transcripts rather than one screenshot, then state plainly how often the payload landed. Ask who wrote the methodology for AI features and when it was last updated, since anything predating 2024 will not mention indirect injection. Check the consultant’s qualifications as well as the company’s accreditations, because an experienced penetration testing companycan still put a junior on your engagement if you do not ask.
Frequently asked questions about AI and application testing
Two questions dominate client conversations on this subject.
Can AI replace a penetration tester?
Not for the findings that matter. Tools have got better at generating payloads and sifting output, but authorisation and business logic flaws depend on knowing what your application is for. No scanner knows that a discount code should not stack four times.
Does an AI feature need separate scoping?
Yes. Add days for it. The tester needs to map every tool and data source the model can reach, which is closer to reviewing an integration than testing a page, and that work does not fit inside the estimate for the rest of the application.

More Stories
Unlock the Future: How AI, Quantum Computing, and Bioengineering Are Redefining Humanity’s Next Chapter
Unlock the Future: How AI, Quantum Computing, and Robotics Are Redefining Humanity’s Next Chapter
Unlock the Future: How AI, Quantum Computing, and the Metaverse Are Redefining Reality