AI red teaming for production GenAI

We find what your pentest and GRC missed.

Cross-tenant data exposure, prompt injection that survives your guardrails, and model supply-chain failures. Human-led adversarial testing against the AI you already have in production, with reproduction steps for every finding.

Two ways in: an automated pass inside 24 hours, or a human-led assessment over one to two weeks. Either way you get findings you can reproduce, not a compliance report.

finding-01 / broken object-level authorisation CONFIRMED EXPLOITED
target enterprise AI sales intelligence platform context active GRC motion, third-party compliance tooling smallest customer $1B+ revenue [01] authenticate as free-trial user ........ ok [02] account IDs are sequential integers .... yes [03] report endpoint ownership check ........ NONE [04] accounts probed ........................ 201 [05] customer PDF reports retrieved ......... 95 [06] cross-tenant intelligence exposed ...... 3.49 MB [07] embedded storage creds, expiry ......... 7 days severity CRITICAL session one total 8 critical or high across 17 vulnerability classes

Published with the client's permission. Read the full case study

95
customer reports retrieved
3.49 MB
cross-tenant intelligence, one request
Active
their GRC tooling
28
findings on another client's MCP surface

A scanner looks for a malformed request. There was nothing malformed here. The endpoint accepted a valid session token and a sequential integer, and simply never asked whether the caller owned the account they were requesting. Authorisation logic that reads as ordinary traffic is the gap we work in.

01  /  The gap

Why your current coverage does not reach this

Nothing below is a criticism of your pentest or your auditor. They are answering different questions, correctly. Neither question is "can this model be talked into acting outside its authority".

What you already run What it answers What it cannot reach
VAPT / pentest Is the request path exploitable Authorisation expressed as an instruction to a model, not as a rule in code
GRC / ISO 42001 / SOC 2 Do you have controls, and evidence of them Whether a control holds under adversarial pressure
AI-SPM scanners Known patterns and signatures at scale Novel chains that require reasoning about your specific architecture
Our own automated pass Our attack library, run end to end in under a day The same ceiling. We run one, we sell one, and we say where it stops
Guardrail products Blocking known-bad prompts Attacks that never look like a bad prompt
02  /  Two ways in

Automated in a day, or human-led in a fortnight

One category, two velocities. They answer different questions, and most clients want both in sequence rather than one instead of the other.

24 hours or less AWS Marketplace

Automated assessment

Answers what is already broken. Known attack classes run end to end with no human in the loop, against your live system. Same attack library our own team works from.

Use it when

  • You need a baseline before a launch, and you need it now
  • A customer questionnaire needs AI testing evidence this week
  • You want regression cover between human assessments
  • Procurement is the bottleneck, so buying on your existing AWS commit is worth more than a new vendor form

What it will not do. It finds what it has patterns for. "Does this endpoint check ownership" is not a pattern, which is why the finding at the top of this page needed a human.

1 to 2 weeks Human-led

Human-led assessment

Answers what can be broken. Architecture-aware adversarial testing that reasons about your specific tenancy, tool wiring and retrieval scope, looking for chains nobody has written a signature for yet.

Use it when

  • The system is multi-tenant, and a boundary failure is material
  • The AI can take actions, not just answer questions
  • A board, an auditor or an enterprise customer is asking
  • The automated pass came back clean and you do not believe it

What you get. Exploit chains with reproduction steps, severity, and a fix path per finding. Every finding published on this site came out of this work.

Automation covers the known-pattern majority cheaply, so human judgment goes where it is actually needed. We will tell you which one you need, including when the answer is the cheap one, and including when it is neither.

04  /  Who we work with

Three readers, one category

Security leadership

You own AI risk you did not choose

Product shipped an assistant, and it reached production without anything resembling an adversarial review. You need to know what is actually reachable before someone else finds out.

Founders and CTOs

An enterprise deal is stuck in security review

Clear the AI section of your customer's security questionnaire with evidence rather than assurances. A focused one-week red team usually unblocks it.

Consultancies and SIs

You need specialist capability you cannot hire

The AI red team you white-label into your engagements. We stay behind your brand and never approach your client directly. How the partner model works

05  /  The alternatives

What you are probably comparing us to

Worth being direct about this, because you are going to ask anyway.

Option What you get Where it falls short
Big 4 AI risk assessment Six to eight weeks, a compliance-style report, a recognisable name on the cover Findings you cannot reproduce, written by people who did not attack the system
Platform AI-SPM scan Dashboards, signatures, breadth, often already inside existing spend Cannot reach anything that needs reasoning about your architecture, and rarely says so
HiltLock, automated Under 24 hours, known attack classes, bought on AWS Marketplace against your existing commit Pattern-bound by design. We tell you that in writing rather than leaving you to find out
HiltLock, human-led One to two weeks, actual exploit chains with reproduction steps, a named human accountable for the work Depth over breadth. We do not do traditional VAPT, and we will tell you when you do not need us
06  /  Quarterly findings report

What we are finding, every quarter

Anonymised patterns from the engagements we ran this quarter: which failures recurred, which controls held, and how long each took to bypass. No vendor pitch, no gated demo call.

Quarterly Findings Report / Excerpt HiltLock

4.2 Authorisation expressed in natural language is not enforced

The most frequent Critical finding this quarter was not a model weakness. It was retrieval scope enforced by instruction rather than by policy, an authorisation boundary described to the model instead of imposed on it.


Observed in 6 of 9 engagements / Median time to first bypass: 2.5 days

Find out what is actually reachable

Twenty minutes, no deck. Describe your AI system and we will tell you whether we are the right people, and what we would go after first.