One arc, three phases
Assess what is reachable today. Monitor it as your system changes. Control the boundaries that should never have depended on a prompt. Most clients start at Assess and stay there for a while, which is usually the right call.
Two velocities, two different questions
The automated pass and the human-led assessment are not a cheap version and a good version. They answer different questions, and running the fast one first is often the cheapest way to find out whether you need the slow one.
Automated assessment
What is already broken? Our attack library run end to end against your live system with no human in the loop. The same library our own team works from, which is why it is worth running even when a human assessment is coming.
How it runs
- Subscribe on AWS Marketplace and run it yourself, or have us run it and hand back the report
- Results the same day in most cases, inside 24 hours in all of them
- Output is a findings report with severity and reproduction steps, not a dashboard you have to interpret
- Schedule it per sprint once you are past the first run
Where it stops. It finds what it has patterns for. A novel chain through your specific tenancy model is not a pattern, and no automation reaches it. We would rather write that here than let you discover it later.
Human-led assessment
What can be broken? Architecture-aware adversarial testing. We read how your system is actually wired, then go after the boundaries that only exist because something told the model to respect them.
What we go after
- Tenant and authorisation boundaries expressed in natural language rather than in policy
- Prompt injection that survives your guardrails, including indirect injection through retrieved content
- Agent and tool-use chains where the model can take an action it should not be able to reach
- Retrieval scope, and what is actually in the index that should not be
- Model supply chain, and shadow AI surfaces nobody registered
What you get. Exploit chains with reproduction steps, severity, and a fix path per finding. Every finding published on this site came out of this work.
Automation covers the known-pattern majority cheaply, so human judgment goes where it is actually needed. That split is deliberate. We will tell you which one you need, including when the answer is the cheap one, and including when it is neither.
What it costs
Both currencies are shown to everyone, everywhere. You should not have to wonder whether the number depends on where you are reading from.
Automated assessment
Per run, standard scope.
- Results inside 24 hours
- Run it yourself, or we run it and hand back the report
- Subscription packages available for scheduled runs
- Procurable against your existing AWS commit
Standard scope covers a single production system. Price moves with complexity and how much of the attack library needs tailoring to your architecture.
Human-led assessment
Scoped per engagement.
- Priced against system complexity, not headcount or hours
- Quarterly re-test subscriptions run at a standing rate
- You get a number in the first conversation, not after a discovery phase
- We will tell you if the automated pass is enough
Multi-tenant systems, agents that can take actions, and anything with a regulator attached sit at the top of the range. A single-surface assistant sits near the bottom.
Automated pricing is per run at standard scope, and the same whether you buy direct or through AWS Marketplace. Human-led engagements are scoped on the first call, because a single-surface assistant and a multi-tenant agent platform are not the same job.
A clean report has a shelf life
You changed the model version. Someone tightened a system prompt. A new tool got wired into the agent. Every one of those can reopen something an assessment closed, and none of them looks like a security change on the way past.
Monitor is two cadences running together. Automated runs on your sprint rhythm catch regression against known patterns. A human re-test each quarter catches what the automation structurally cannot, and re-reads the architecture for what changed shape.
We will not claim a full red team on every commit. That is neither feasible nor useful, and anyone selling it is selling you the automated tier with better marketing.
Stop asking the model to enforce its own limits
Most of what we find comes down to one pattern: an authorisation boundary that was described to a model instead of imposed on it. Control is pre-execution policy for those boundaries, evaluated before the model or the agent acts.
After evidence, not before
We deploy this where an assessment showed a boundary that cannot be fixed in the application. Not as a default, and not as a platform you have to adopt to work with us.
Policy before action
Identity and scope checks on retrieval and tool calls, so the answer to "may this request reach this data" stops depending on how the request was phrased.
A finding, not a product
Unregistered AI surfaces show up during Assess as a finding category. Where they need governing rather than just documenting, that governance happens here.
What we do not do
Worth stating plainly, because it is usually the second question and because a specialist who will do anything is not a specialist.
| Not us | Why | What we do instead |
|---|---|---|
| Traditional VAPT and network pentest | A price-compressed commodity, and doing it badly would cost us the specialism | We prime on the AI work and subcontract the commodity scope to people who are good at it |
| Security operations, monitoring, SOC | A staffing business wearing a security badge | We hand findings to whoever runs your operations |
| Certification or attestation | We are not an auditor, and pretending otherwise would be worth less than the testing | We produce the evidence your auditor and your customers ask for |
Start with the cheap question
Run the automated pass, see what comes back, then decide whether the human assessment is worth it. Twenty minutes on a call is enough for us to tell you which way to go.