Voice agents, attacked over a real call
A voice agent that holds up in a text harness can still be talked out of its authority on a live call. We place the call, over SIP and PSTN or over WebRTC, and go after the agent on the channel it actually runs on.
Voice is a surface in our existing engine, not a separate product. It is priced and scoped inside Phase 01 like every other surface, and it can run in the same engagement as your text testing.
Illustrative format, not a client finding. Every real finding we publish is at Evidence.
What changes when the channel is audio
Testing the model behind a voice agent through its text API tells you about the model. It does not tell you about the agent, because most of what makes a voice agent exploitable lives in the parts that only exist on a call.
| What a text harness reaches | What only a real call reaches |
|---|---|
| Whether the model refuses a written request | Whether the refusal survives being talked over before it finishes |
| Filters matching on written text | The same payload spelled out, paced, or framed phonetically through a speech-to-text layer |
| A single request and response | Authority built up across a whole conversation, where no single turn looks like an attack |
| An authenticated API session you were given | Whether a session can be minted without authentication, and whether its scope and expiry can be tampered with |
| The agent's own answers | What the agent tells the human it transfers you to, and what it hands over with the call |
Very few firms will test a voice agent over live telephony, because it requires carrier plumbing and real-time audio tooling rather than an HTTP client. That is the whole reason this surface stays untested in most deployments.
What we go after on a voice agent
Six classes, all of them scored on the call rather than on a transcript we generated ourselves.
Cumulative instruction extraction
What the agent gives up across an entire conversation rather than in any one reply. Scored over the whole call, because a per-turn judgement misses the disclosure that took nine turns to assemble.
Voice-native evasion
Paced delivery, spelling out, and phonetic framing. Payloads that a text filter would catch and a speech-to-text layer reassembles into something it never saw.
Interruption and barge-in
Talking over the agent mid-sentence to cut off its safety preamble. A control that is only ever spoken can be removed by not letting it finish.
Session control plane
Unauthenticated session minting, and JWT scope and expiry tampering. This is ordinary authorisation work, and on voice stacks it is routinely the most serious thing on the call.
Identity and authority pressure
Claiming to be a supervisor, an engineer, or a verified customer. Whether the agent's idea of who it is talking to can be set by the caller.
Human-transfer boundary
Where the agent hands off to a person. What it passes along, what it asserts about the caller, and whether the handoff can be used to launder a request the agent already refused.
Delivery is verified, not assumed. If the target could not have heard a payload, we report it as undelivered. We never report it as a pass. An attack that failed to arrive tells you nothing about your agent, and counting it as a pass would inflate every number on the report.
How we reach your agent
SIP and PSTN
A real phone call to the number your customers call. Carrier path, codec and latency included, because all three change what an attack can do.
WebRTC and WebSocket audio
In-app and in-browser voice, and embedded agents that never touch a phone number. Same classes, different transport.
Whatever you built it on
Twilio, Telnyx, LiveKit, OpenAI Realtime, ElevenLabs, Deepgram, or your own protocol. We test the agent, not the vendor.
One engagement, both channels
Most organisations running a voice agent are running a text interface to the same system. Testing them together is worth more than testing them separately, because the interesting finding is usually the difference: the guardrail that holds on text and does not hold on audio. That comparison only exists if one engine runs both, which is why voice is a surface in our engine rather than a separate product.
Speeds
- A focused run against a named class or two, roughly one to two hours
- A full-spectrum pass across all six classes, inside 12 to 24 hours
- Human-led, where the voice surface is one part of an architecture-aware assessment over one to two weeks
The first two are the automated pass described on the program page, pointed at the voice surface.
What we cap, and what we keep
Two questions come up every time, and both deserve an answer before you are on a call with us.
Your bill stays predictable
Every run carries hard caps on the number of calls, the number of turns per call, and total duration. Voice testing consumes your telephony minutes and your model inference, so an uncapped run is a bill nobody agreed to. The caps are set before the run starts and the run stops at them.
Recordings, and how long we hold them
- Audio attached to a finding is retained 90 days, so you can hear the exploit rather than take our word for it
- Audio not attached to a finding expires automatically after 7 days
- Encrypted at rest, and reachable only through an authenticated request scoped to your own run
Where this stands today. The first voice engagements are now published on the red team platform: a voice agent that gave up its operating instructions across four spoken turns, and a widget whose session endpoint minted billable sessions unauthenticated. Both ship with the same evidence standard as everything else - the call recording and transcript, or the replayable request/response pair - and are published with client permission, on the same terms as all findings.
Point it at your voice agent
Tell us the number or the endpoint, what the agent is allowed to do, and who it hands off to. Twenty minutes is enough for us to say what we would go after first.