AI Red Teaming: Capability Is Not Competence
Reflecting on the recent progress AI labs have announced in cybersecurity, one thing is becoming clear: the labs are successfully shaping the narrative, but the reality is more prosaic.
They have built LLMs that can sometimes find and exploit vulnerabilities. That is real progress. It matters. But it is not yet autonomous cybersecurity in any meaningful operational sense.
Capability is not the same as competence
Current models remain weak at genuinely complex work and long-horizon agentic tasks. Context degrades. Assumptions drift. Unexpected conditions appear. The model loses track of what it is doing, or optimises for the next plausible step rather than the right one.
That failure mode is easy to miss in a demonstration, because a demonstration is short, the environment is known, and someone has already decided in advance what success looks like. Real environments are full of ambiguity, brittle dependencies, incomplete information, and consequences that cannot be undone with a new prompt.
None of this means the capability is fake. It is not. We have written elsewhere about what agentic security tooling actually achieved in 2026, and those results are not marketing. A model that reasons about a compression algorithm until it works out the precise conditions under which the algorithm’s own assumptions break is doing analysis, not pattern matching.
But finding a bug in a library is a bounded task with a checkable answer. An engagement is not.
What “AI red teaming” often looks like in practice
So far, much of what is presented as “AI red teaming” looks like an LLM from one company, unrestrained by the inadequate controls of a second company’s testbed, hacking systems operated by a third company.
A reputable penetration-testing firm or red team would never do that. Neither would it run a PR campaign around such a screw-up.
Before a competent team touches anything, it establishes authorisation, defines scope, understands the environment, coordinates safeguards, protects evidence and data, and knows when to stop. None of that is bureaucratic overhead. It is the entire difference between a test and an incident. It is also most of what threat-led penetration testing formalises, and most of what a client is actually paying for.
The model does none of it.
The model has no duty of care
It does not possess professional judgment. It has no duty of care towards a client, no understanding of proportionality, and no internal sense of what should remain untouched.
If we fall into the sin of anthropomorphising it, the best metaphor is not a wise assistant or a trusted adviser. It is a desperate hostage with Stockholm syndrome, eager to help whoever happens to be holding the prompt.
That may make it useful. It does not make it safe.
Who the standards are actually protecting
The industry’s mission has always been to protect clients. Right now, that increasingly means protecting them from decisions that are made too quickly, on the back of capabilities that have not been adequately tested or understood.
Many red teams are already experimenting with AI, and some have pushed it far enough to understand its practical limits. Clients usually have not. They are exposed to a loud mixture of AI lab claims, startup marketing, and demonstrations that confuse isolated capability with dependable operational competence.
Too many vendors sell certainty when they have barely established repeatability.
That word carries more weight than it appears to. Repeatability is the minimum evidence that a result belongs to a method rather than to a lucky run. A single impressive finding is an anecdote. A finding that a team can reproduce, explain, and defend under scrutiny is a service. Telling one from the other is precisely the work a buyer does when choosing a security provider.
Judgment is the layer nobody has automated
AI labs know how to build capable models. But capable models are not ethical operators. What they lack, for now, is the judgment to decide what to do next, what not to do, and when to stop.
Attempts to compensate for this through deterministic harnessing will either fail or reduce AI productivity below an acceptable level. Constrain an agent tightly enough to guarantee it never crosses a line, and you have rebuilt a scanner. Loosen it enough to be genuinely useful, and the guarantee is gone. No setting on that dial produces judgment, because judgment is not a constraint. It is a decision made in context by somebody who is accountable for it.
This is where experts must step in. Red-teamers, pentesters, security architects, lawyers, and client leaders need to apply the ethical judgment that algorithms do not have.
The task is not to resist the technology or pretend it changes nothing. It is to introduce it deliberately: with clear authorisation, constrained scope, human oversight, tested workflows, auditable evidence, and accountability for the outcome. That is the same discipline a serious LLM penetration test demands of the systems it examines. It applies just as much to the tools we use to examine them.
Six questions worth asking any vendor selling AI-driven offensive security
- Who authorised the testing, and does that authorisation cover everything the agent can actually reach?
- What is the agent permitted to do, and what stops it at the boundary? Name the control, not the intention.
- Which parts of the work were autonomous, and which were supervised? A blended answer is a non-answer.
- Can you reproduce this finding, or did it happen once?
- What evidence exists that a qualified human reviewed the result before it reached me?
- Who is accountable if the agent causes damage, and what does the contract say about it?
A vendor who has done this work seriously will answer all six without discomfort. A vendor who is selling a demonstration will reframe the questions.
The question that actually matters
The question is not whether AI will become part of offensive security work. It already is, in tooling, in triage, and in the parts of the job that are bounded enough for it to be dependable.
The question is whether we let it erode the professional standards that make offensive security legitimate in the first place.
BSG runs penetration tests the way this article describes — explicit authorisation, agreed scope, human testers accountable for every finding, and evidence you can reproduce and defend. We use AI where it is dependable and say plainly where it is not.
Talk to a tester →