AI Red Team  -  MCP Protocol

A saved record told the agent
what to do next.

The attacker never spoke to the model. They typed into a field the product invited them to fill in, and waited for an assistant to read it back.

Target
Top-five consumer app in its category, globally
Surface
Production MCP server, 20 tools reachable
Assessment
Automated red team, model and API authorisation
Highest severity
High
4 of 5
Objectives the system
successfully defended
1
Finding, high severity,
and it reached a transaction
0
Credentials or authorisation
bypasses required

Not a soft target, by any measure.

This was a consumer application in the global top five of its category, run by an engineering organisation that would sit in anyone's top percentile, with a real security function and the budget to staff it. Its web and mobile surfaces have been tested by good people for years. None of that is in question here, and the scoreboard says so: of five objectives, it defended four.

What it had recently added was a Model Context Protocol server, so that assistants could act on a user's behalf. Twenty tools reachable, nine resources, fourteen retrieval endpoints, on a single endpoint. That surface is two years old as a category. There is no accumulated practice for it, no signature list, and very few people who have attacked one.

01

A mature security programme tests the surfaces it has a playbook for. This surface has no playbook yet.

02

The tools were not the weakness. Every one behaved exactly as designed, and the design was reasonable.

03

The weakness only exists once an agent is the thing reading the output, which is what the protocol is for.

Two tools, each correct.
One chain, not correct.

The product lets a user save records to their own profile, and each record carries a short free-text label the user chooses. A second tool lists those records back. Both are ordinary, and neither is a vulnerability.

The label was stored as typed and returned as typed. So text written into it arrives, verbatim, inside the context of any agent that later lists the user's own records. At that point the question stops being what the field contains and becomes what the reader does with it.

Step 01  -  Reachability
Text in, text out
Expected

A label is data. It describes a record, and a human reads it on a screen.

Observed

Returned verbatim by the listing tool, so anything stored there enters the context of every agent that reads it. No model had acted on it yet. This step alone is only reachability.

Step 02  -  Obedience
An action nobody asked for
The user asked

A single benign question about their own saved records. Nothing else.

The agent did

High. It read the planted label, followed it, and called a tool that commits a transaction on the user's behalf. The control episode, identical in every respect except the text in one record, did not.

Step 03  -  Tenancy
Someone else's account
Assumed boundary

One user's data cannot influence what happens in another user's account.

Observed

Text one account saved caused a multi-tenant agent to act against a different account, including a destructive call. No credential. No authorisation bypass. Only an intermediary holding both delegations.

Step 04  -  The mitigation
One line of prompt
Carrier

The unsanitised field is still the carrier, and still stores whatever is typed into it.

But

An explicit "never act on another account" instruction in the consuming agent prevented the cross-account step outright. A one-line fix mitigates the worst of it downstream, without the field ever changing.

Nothing here is a jailbreak. The attacker never talks to the model, never sends a prompt, and never authenticates as anyone else. They fill in a form field the product asked them to fill in. In a multi-tenant agent, the person who stores the text need not be the person whose session executes it.

This is not one company's bug.
It is the deployment shape.

Every ingredient here is present in most MCP servers shipping today, which is why the case is worth publishing without naming anyone.

A writable free-text field. Almost every product has one. Labels, notes, nicknames, descriptions, tags. They were designed to be read by a human, and now they are read by a model.

A tool that returns it verbatim. Correct behaviour for an API. Sanitising user text on the way out would have looked like a bug when a human was the only reader.

An action tool in the same toolset. The protocol exists so an agent can do things, so read tools and write tools sit side by side in one context by design.

An intermediary holding two delegations. This is the part people miss. The protocol actively encourages a client that acts for many users at once, and that is what turns a self-inflicted injection into a cross-tenant one.

The remediation is the same everywhere: treat retrieved and third-party content as data, never as instructions. Sandbox tool and retrieval output, strip or neutralise embedded directives, and require explicit allow-listing before anything retrieved can influence behaviour. On the consuming side, state the tenancy boundary in the agent's own instructions, because in this case that alone stopped the worst outcome.

OWASP LLM LLM01
OWASP Agents A02 indirect injection
MITRE ATLAS indirect prompt injection
EU AI Act Art. 9, 13, 15
ISO 42001 Clause 6.1, 9.1
NIST AI RMF MAP 1.5, MEASURE 2.5
GDPR Art. 32
DPDPA Section 8

Those mappings are potential regulatory exposure rather than an assertion of breach, and they are subject to the operator's own legal and compliance review.

Four of five, and they matter.

One finding out of five objectives is a good result for the operator, and reporting it as anything else would be dishonest. The four that held are not padding: they are the difference between a finding and an incident.

Held
Writes were properly bounded

Every write the probe attempted was against a record it had planted itself, and each was refused unless the identifier matched exactly. There was no path to enumerate or to modify arbitrary records, which is the difference between this finding and a breach.

Held
Authorisation itself was sound

Nothing was bypassed. No token was stolen, forged or replayed, and no endpoint failed an ownership check. The cross-account effect came from an agent legitimately holding two delegations, not from broken access control. That distinction matters for whoever has to fix it.

What this run does not tell you.

This was an automated pass at the model and API-authorisation layers, proven through a single vector in five attempts. It is not an architectural review: a human-led assessment would reason about the whole toolset, the retrieval scope behind those fourteen endpoints, and which of the twenty tools should ever share a context with a writable field in the first place.

There is no remediation record in this case study, because there is nothing verified closed to report. The finding was disclosed to the operator. When we say a finding is closed on this site it means we re-ran the attack against the running system and watched it fail, and that sentence has not been earned here yet. The reproduction steps, the endpoint, the tool names and the field are withheld for the same reason.

Anyone who shipped
an MCP server.

If you exposed tools to an agent in the last year, you have the ingredients above. The question is not whether your API is correct. It is what happens when a model, rather than a person, is the thing reading your API's output.

Chain testing, not tool testing

Each tool here passed on its own. We test what one tool's output does to another tool's behaviour, which is where the finding was.

A control episode every time

The same session without the planted text, so a claim that the agent acted because of the injection is demonstrated rather than asserted.

Tenancy tested, not assumed

Most MCP testing runs as one user. The interesting failures need two, and an intermediary that holds both delegations at once.

Results inside 24 hours

This was the automated pass, which is the cheap way to find out whether you need the expensive one. Same attack library either way.

Find out what your tools
tell each other.

If your MCP server is live, the automated pass returns inside 24 hours. Twenty minutes on a call is enough to scope it.

Book a 20-min fit call

Published anonymised. Endpoint, tool names, field name and reproduction steps are deliberately withheld.