AI red teaming: one arc, three phases
Assess what is reachable today. Monitor it as your system changes. Control the boundaries that should never have depended on a prompt. Most clients start at Assess and stay there for a while, which is usually the right call.
The whole program on one screen
Three phases, in the order they normally happen. You can stop after any one of them, and most clients stop after the first for a good while.
| Phase | The question it answers | When it runs | How it is priced |
|---|---|---|---|
| 01 / Assess | What is reachable in this system today? | Once, at one of two speeds: 24 hours automated, or one to two weeks human-led | Automated is list-priced per run. Human-led is scoped on the first call |
| 02 / Monitor | Has anything we closed reopened since you shipped? | Continuing: automated on your sprint cadence, human re-test each quarter | Subscription, against the cadence you pick and the scope set at Assess |
| 03 / Control | Which boundaries should stop depending on the model behaving? | Only where a finding from Assess or Monitor shows it is needed | Scoped per deployment. We will tell you when no finding justifies it |
Assess is the only phase most clients buy to begin with, and the only one we will recommend without a finding behind it. Pricing for all three is at the foot of this page.
Two velocities, two different questions
The automated pass and the human-led assessment are not a cheap version and a good version. They answer different questions, and running the fast one first is often the cheapest way to find out whether you need the slow one.
Automated assessment
What is already broken? Our attack library - 19 families plus an always-on credential-reuse pivot, each mapped to OWASP and MITRE ATLAS - run end to end against your live system with no human in the loop. The same attack library the human assessments run on, which is why it is worth running even when a human assessment is coming.
How it runs
- Subscribe on AWS Marketplace and run it yourself, or have us run it and hand back the report
- Results the same day in most cases, inside 24 hours in all of them
- Output is a findings report with severity and reproduction steps, not a dashboard you have to interpret
- Schedule it per sprint once you are past the first run
Where it stops. It finds what it has patterns for. A novel chain through your tenancy model is not a pattern, so no automation reaches it.
Human-led assessment
What can be broken? Architecture-aware adversarial testing across every surface the system exposes: conversational AI, RAG and retrieval, APIs and authorisation, tools and MCP servers, agentic workflows, and voice. We read how your system is actually wired, then go after the boundaries that only exist because something told the model to respect them.
What we go after
- Tenant and authorisation boundaries expressed in natural language rather than in policy
- Prompt injection that survives your guardrails, including indirect injection through retrieved content
- Agent and tool-use chains where the model can take an action it should not be able to reach
- Retrieval scope, and what is actually in the index that should not be
- Model supply chain, and shadow AI surfaces nobody registered
- Voice agents, attacked over a real telephony or WebRTC call rather than through a text API
What you get. Exploit chains with reproduction steps, severity, and a fix path per finding. Every finding published on this site came out of this work.
Automation covers the known-pattern majority cheaply, so human judgment goes where it is actually needed. That split is deliberate. We will tell you which one you need, including when the answer is the cheap one, and including when it is neither.
Voice agents are a surface here too
If the system answers a phone or takes microphone input, we attack it on that channel. We place the call. We talk over the safety preamble to see whether the control survives being cut off. We spell payloads out to get past filters that only read text. We push on who the agent thinks it is talking to. Session minting and token scope get tested as on any other surface.
The comparison is the valuable part. A guardrail can hold on your chat interface and fail on audio, and you only see that if one engagement covers both. Voice lives in the same engine here, not in a separate product.
Run shapes
- Focused run against one or two classes, roughly one to two hours
- Full-spectrum voice pass, inside 12 to 24 hours
- Part of a human-led assessment, where the voice surface is one piece of the architecture
Before you ask
- Hard caps on calls, turns and duration, agreed before the run, so your telephony bill stays predictable
- Call audio behind a finding is kept 90 days; audio without a finding expires after 7
- If a payload could not have been heard, it is reported as undelivered and never as a pass
The first voice findings are now published on the red team platform: a voice agent that gave up its operating instructions across four spoken turns, and a widget whose session endpoint minted billable sessions unauthenticated.
A clean report has a shelf life
You changed the model version. Someone tightened a system prompt. A new tool got wired into the agent. Every one of those can reopen something an assessment closed, and none of them looks like a security change on the way past.
Monitor is two cadences running together. Automated runs on your sprint rhythm catch regression against known patterns. A human re-test each quarter catches what the automation structurally cannot, and re-reads the architecture for what changed shape.
We will not claim a full red team on every commit. That is neither feasible nor useful, and anyone selling it is selling you the automated tier with better marketing.
Stop asking the model to enforce its own limits
Some findings you fix in code or architecture. A failure that comes from the model's own generative behaviour cannot be fixed that way, it needs enforcement outside the model, before it acts. That is Control. We recommend it only where a finding needs it, never by default.
Only where the finding needs it
Control is for boundaries that cannot live in the application. Where the fix belongs in your code, the report says so instead.
Policy before action
A gateway in front of your AI traffic runs identity and scope checks on retrieval and tool calls. Whether a request can reach given data stops depending on how that request was phrased. Observe mode logs the decision without touching your traffic; enforce mode applies it.
A finding, not a product
Unregistered AI surfaces show up during Assess as a finding category. Where they need governing rather than just documenting, that governance happens here.
Latency, fail behaviour, hosting and who owns the policy. Written for whoever has to approve something sitting in your traffic path.
What it costs
Both currencies are shown to everyone, everywhere. You should not have to wonder whether the number depends on where you are reading from.
Automated assessment
Per run, standard scope.
- Results inside 24 hours
- Run it yourself, or we run it and hand back the report
- Scheduled runs are priced under Monitor, alongside this card
- Procurable against your existing AWS commit
Standard scope covers a single production system. Price moves with complexity and how much of the attack library needs tailoring to your architecture.
Human-led assessment
Scoped per engagement.
- Priced against system complexity, not headcount or hours
- Standalone, or as the baseline a Monitor subscription then runs against
- You get a number in the first conversation, not after a discovery phase
- We will tell you if the automated pass is enough
Multi-tenant systems, agents that can take actions, and anything with a regulator attached sit at the top of the range. A single-surface assistant sits near the bottom.
Scheduled re-testing
Subscription, on your cadence.
- Automated runs on your sprint rhythm, priced per run against the package you pick
- Human re-test each quarter at a standing rate, agreed before the first one
- Scope carries over from Assess, so there is no second discovery exercise
- Cancellable at the end of any quarter. It is a subscription, not a lock-in
The number moves with how often you want the automated runs and whether the quarterly human re-test is in. Systems that change shape every sprint cost more to keep honest than ones that do not.
Runtime enforcement
Scoped per deployment.
- Priced against the surfaces enforced and where the gateway runs
- Quoted only after a finding shows instruction-level control is not holding
- Self-hosted in your private cloud or consumed as SaaS, which changes the number
- Observe-mode pilot first, so you buy enforcement against your own traffic
Not the phase we lead with. It exists because some findings cannot be closed in application code, and it is quoted per deployment because no two policy layers sit in the same place. How the gateway works.
Automated pricing is per run at standard scope, and the same whether you buy direct or through AWS Marketplace. Everything else here is scoped rather than list-priced. A single-surface assistant and a multi-tenant agent platform are not the same job, and one flat price would overcharge the smaller one. You get a number in the first conversation either way.
What we do not do
Usually the second question we get. A specialist who will do anything is not a specialist.
| Not us | Why | What we do instead |
|---|---|---|
| Traditional VAPT and network pentest | A price-compressed commodity, and doing it badly would cost us the specialism | We prime on the AI work and subcontract the commodity scope to people who are good at it |
| Security operations, monitoring, SOC | A staffing business wearing a security badge | We hand findings to whoever runs your operations |
| Certification or attestation | We are not an auditor, and pretending otherwise would be worth less than the testing | We produce the evidence your auditor and your customers ask for |
Start with the cheap question
Run the automated pass, see what comes back, then decide whether the human assessment is worth it. Twenty minutes on a call is enough for us to tell you which way to go.