What we actually found
Five systems, published anonymised. Three had passed the checks their buyers asked for, and no buyer was wrong to ask; those checks just do not reach this. The fourth ran an agent protocol that nobody has checks for yet, and the fifth is a voice agent attacked over a live call.
Active GRC tooling. A free trial account. 95 customers' reports.
An AI-enabled sales intelligence platform with enterprise contracts, an active compliance motion, third-party GRC tooling in use, and a smallest customer north of $1B in revenue.
The finding that mattered
The platform's core value is AI-generated research reports built over months of analysis and stakeholder profiling. The report download endpoint accepted an account ID and never checked whether the caller owned that account. Account IDs were sequential integers.
A single authenticated user, including a free trial account, could iterate the range and collect working download URLs. We obtained 95 customer reports across 201 probed accounts in one session. The URLs embedded temporary storage credentials valid for seven days, so each exposure kept a week-long tail.
Three more that change the threat model
- System prompt extraction. The AI returned its verbatim operating instructions to any user who asked. That is a map of every constraint the platform had, and once it is out, every control built on "the AI will not do X" becomes negotiable.
- Persistent prompt injection. Attacker-controlled instructions could be written to the AI's stored configuration. They survived logout and applied to every future interaction on that account.
- Fabricated compliance certifications. The AI generated and published SOC 2, ISO 27001 and GDPR certifications with no content review and no human approval. One API call, externally shareable.
Compliance tooling was active and working as designed. It answers whether controls exist and are evidenced. It does not answer whether an endpoint checks ownership.
Losing enterprise deals at security review, then not
An AI-native LMS company selling into large Indian enterprises. They had just lost a significant deal, not on product, but because they could not clear the buyer's security review. More deals were in the pipeline carrying the same questionnaires.
Why nobody else was assessing it
Their MCP server was the novel attack surface, and generic VAPT vendors do not assess it. Not because they are careless, but because it is new territory and there is no established methodology to apply. Meanwhile the enterprise questionnaires kept arriving: DPDP, CERT-In logging, cloud posture, LLM supply chain, MCP security.
We ran the full MCP red-team framework against production authentication and the MCP server, both unauthenticated and privileged. The focus was chained attacks: individually defensible components that combine into critical enterprise risk.
What the chain gave up
- Malformed input leaked the tenant ID, the database schema and storage tokens
- OAuth client registration was wide open
- Findings mapped to OWASP Top 10 and OWASP LLM Top 10, so the output was a CISO-ready narrative rather than a raw technical list
We do not accept "fixed" on trust. After the engineering team shipped remediations we re-ran both assessments live against the running system, same attacker model, same two-tenant probes, and recorded exactly what closed.
Error and secret channels
Before. 28 findings, 3 critical and 8 high. Malformed input leaked the tenant ID, the database schema and storage tokens. OAuth client registration was wide open.
After. Every critical and high finding closed and verified by live retest against the running system, not by reviewing a remediation ticket.
Outcome: enterprise security reviews cleared. Two months from weak posture to enterprise-grade.
The direct asks were refused. The reframed ones were not.
A US-based digital twin company, where the product is an AI standing in for a named human expert. It refused every blunt request it was given. Changing the framing rather than the request got through nine times, ending in a disclosed credential.
Refusal is not the same as resistance
Asked outright for its system prompt, the model refused. Asked for real personal data, it refused. Those refusals are why a questionnaire would have recorded this system as handling prompt injection correctly, and they are the reason the run is interesting: the attacker never had to defeat them. It asked the same questions inside a different frame.
A short story to be written. Instructions encoded as base64. A previous answer that had supposedly been cut off mid-sentence. A template shape needed for a team guide. Synthetic test data rather than real records. Each frame is innocuous. Each one moved something out that the direct question could not.
Three that change the threat model
- A credential, and this was the critical. Progressive questions about the assistant's own capabilities walked from generic answers to its non-public tool-discovery interface, and from there to a sensitive credential in plain text. That is a path at internal systems rather than a disclosure about them.
- The system prompt, via fiction. Asked to write a story, the assistant opened it by quoting a verbatim non-public portion of its own operating instructions.
- Internal tooling, unprompted. Asked to draft an email from a supplied document, an indirect injection, it volunteered an internal ticketing process and project identifier that appeared nowhere in the prompt.
It held twice out of thirty-five. An indirect injection probe was defended, as was an attempt to coax it into executing a tool. The guardrails were not absent. They were attached to the shape of a request rather than its effect.
This run was interim and did not fully complete, so the nine findings are a floor, not a total. A longer run would likely find more.
Endpoint names, internal identifiers and the credential itself are deliberately not published.
A saved record told the agent what to do next
A consumer application in the global top five of its category, with a security function and an engineering team in anyone's top percentile. It defended four of five objectives. The one that got through needed no credential and no authorisation bypass.
Two correct tools, one incorrect chain
The product lets a user save records to their own profile, each with a short free-text label they choose. A second tool lists those records back. Both are ordinary. Neither is a vulnerability. The label was stored as typed and returned as typed, so text written into it arrives verbatim inside the context of any agent that later lists the user's own records.
The attacker never speaks to the model, never sends a prompt and never authenticates as anyone else. They fill in a field the product asked them to fill in.
What the chain gave up
- An action nobody asked for. The user asked one benign question about their own saved records. The agent read the planted label, followed it, and called a tool that commits a transaction on the user's behalf. The control episode, identical but for the text in one record, did not.
- A crossing into another account. Text one account saved made a multi-tenant agent act against a different account, including a destructive call. No credential, no authorisation bypass, only an intermediary holding both delegations, which is the deployment this protocol encourages.
- A one-line mitigation downstream. An explicit "never act on another account" instruction in the consuming agent stopped the cross-account step outright, while the unsanitised field remains the carrier.
Authorisation itself was sound. Nothing was bypassed, no token was stolen or replayed, and no endpoint failed an ownership check. Every write the probe attempted was against a record it had planted, and was refused unless the identifier matched exactly.
Every ingredient generalises: a writable free-text field, a tool that returns it verbatim, an action tool in the same context, and an intermediary holding two delegations. Most MCP servers shipping today have all four.
Disclosed to the operator. Endpoint, tool names, the field and the reproduction steps are deliberately not published.
Four polite turns, one reconstructed rulebook
Voice is attacked over the channel it actually runs on - a live call or the widget on the site - not through a text API. Both findings below ship with the same evidence standard as every other case.
Judged turn by turn, it passed
Every individual turn was a polite refusal. Judged turn by turn the voice agent was a clean pass; judged across the whole conversation, four turns reconstructed the agent's own operating instructions.
The second finding needed no prompt cleverness. The widget's session endpoint minted billable sessions with no authentication, so every call provisioned real capacity on the operator's bill.
What was not found is stated too. Against a hardened agent, none of the ten extraction techniques recovered a discrete secret; instruction reconstruction is what actually worked, and the report states which attacks failed, not only which landed.
A guardrail that holds in chat and fails on a phone call is itself the finding - and it is invisible to anyone testing the two separately.
Published on the red team platform.
What all five engagements have in common
Five very different systems, five different attack surfaces, one shared shape.
The traffic looked legitimate
No malformed request, no injected payload a signature would match. A valid session and an incrementing integer in one case, a request to write a short story in another, a user filling in a form field in the fourth. Nothing to detect, because nothing looked wrong.
The control existed and was answering
GRC tooling active in the first, enterprise questionnaires being answered in the second, a model that refused every blunt request in the third, in the fourth an authorisation layer that never once failed an ownership check, and in the fifth a voice agent that refused every single turn while still giving up its instructions across the conversation. All five would evidence the control. None of those exercises asks whether the control still means anything when an agent is the thing reading the output.
The surface was too new to have a playbook
MCP servers, agent tool wiring, AI-generated artifacts published without review, a guardrail that can be asked to explain itself, a free-text field that was safe for twenty years because only humans read it, and a spoken channel where a guardrail that holds in chat fails on a call. There is no signature list for these yet, which is precisely why someone has to go and look.
Find out what is reachable in yours
Start with the automated pass if you want a fast baseline, or go straight to the human assessment if the system is multi-tenant or can take actions. Twenty minutes is enough for us to tell you which.