About

Not a security company that learned AI

Built by product and engineering people from large-scale ML, personalisation and identity systems, not by a security team that added AI later. That order is the point: the architecture gets read before it gets attacked.

The name

A blade is only useful if you can hold it

AI is a blade. It is genuinely sharp, which is the entire point, and the companies winning with it are the ones swinging hardest. Nobody sensible responds to a sharp tool by refusing to pick it up.

The hilt is the part that lets you hold it. The crossguard is what stops your hand sliding onto the edge while you are busy doing the thing you picked the blade up to do. It does not make the blade less sharp. It makes it yours to use.

That is the job. We are not here to slow your AI programme down or to tell you the risk means you should not ship. We are here so you can move fast without cutting yourself, accidentally or because somebody else made you.

the mark NOTES
| |______ | | the vertical the blade, in steel the crossbar the guard, in ember also the turnstile operator, written |- in logic, and read as "proves" which is the only claim we make: here is the finding, here are the steps, reproduce it yourself

Ember always marks the part we contribute. Same rule in the wordmark, where it sits on Lock.

The difference

A different question

Most of the industry asks whether a prompt is safe. That is a content question, and it is the one automation is already good at answering.

The question is whether this action, by this identity, on this data, right now, should be allowed. That is an authorisation question, and nearly every published finding came from asking it.

The 95 customer reports a free-trial account walked away with was not a prompt problem. Nothing was phrased cleverly. There was no jailbreak. The system simply never checked who was asking.

the two questions WHY IT MATTERS
the industry asks "Is this prompt safe?" content filtering, signatures, guardrails automation handles this well the question "Should this action, by this identity, on this data, right now, be allowed?" authorisation, tenancy, tool scope needs a human who read the architecture
Tenets

What we hold to when it costs us something

Anyone can list values. These are the ones that have actually cost revenue, the only test that means anything.

01

Evidence over assertion

Every finding ships with reproduction steps. If you cannot reproduce it without us in the room, we have not finished writing it. Severity we assert is worth nothing; severity you can verify is worth acting on.

02

We say where automation stops

Our automated pass is genuinely useful and genuinely limited, and we publish both halves of that. A vendor who will not tell you the ceiling of their own tool is selling you the ceiling as the whole building.

03

Right-sell, not upsell

You will hear us decline revenue. If the cheap automated pass covers your situation, that is what we quote. If neither suits you, we will tell you.

04

Depth over breadth, on purpose

No traditional VAPT, no security operations, no certification. A specialist who will do anything is not a specialist, and the moment we take on commodity work is the moment we stop being worth calling for this.

05

Retest, never take the ticket's word

"Fixed" is a claim until it is re-run against the live system. We re-attack with the same model and record what actually closed. Sometimes the answer is that it did not.

06

The findings are yours

Written so your engineers can act without us, and so your buyers and auditors can read them without translation. No portal you have to keep paying for to see your own results.

Continuity

"You are a small team. What if you disappear?"

Usually the third question we get, and a fair one. Headcount is not the reassurance you actually want, so here is what protects you instead.

01

The method is not in one head

The attack library and testing sequence are written down, not improvised. Findings ship with reproduction steps and a fix path, so your team can act on them with or without us in the room.

02

Findings outlive the engagement

Every finding names the affected components, so nothing needs interpreting by whoever wrote it. Stop working with HiltLock tomorrow and the reports still stand on their own.

03

Escalation in writing

Response times and the escalation route sit in the engagement terms rather than being implied. Anything found that cannot wait for the report reaches you the same day.

Ask us about continuity rather than assuming. A large firm with a rotating bench and a junior on your account is a continuity risk too, just a better-dressed one.

Provenance

Where the instincts come from

Builders' background, over a decade in large-scale ML, personalisation, identity and data-governance systems. The authorisation instinct behind most findings comes from having built the systems that get this wrong.

Amazon Myntra Meesho Zepto

Where the background was built, not a customer list. Client logos do not go on this site without permission, and most prefer not to advertise that they needed us.

HiltLock is operated by CalmSparks Tech Pvt. Ltd. Previously in market as Trampolyne AI. Same methodology, sharper focus.

Common questions

Questions people ask before the first call

What is HiltLock?

HiltLock is an AI red teaming firm. We attack generative AI systems the way an adversary would, before launch or after. Every finding comes with reproduction steps, so your engineers can confirm it and close it. Two speeds: an automated pass inside 24 hours, and a human-led assessment for systems that are multi-tenant or can take actions. Where a finding cannot be closed in your code, we also supply policy enforcement outside the model.

What does the name HiltLock mean?

Two words. Hilt, the handle and crossguard of a blade, the part that lets you hold something sharp without cutting yourself. Lock, what holds once the finding is closed. AI is the blade, and we are not asking anyone to put it down.

If you publish findings, will ours end up on your site?

Not unless you agree to it in writing. Every case on the evidence page is anonymised, and client logos do not go on this site without permission. The default is that an engagement leaves no public trace of who it was.

How is this different from a penetration test?

A penetration test asks whether your infrastructure and application code can be broken. It is good at that. AI red teaming asks whether a legitimately authorised model, fed untrusted input it was designed to read, can be made to take an action your policy would never have allowed. Neither test replaces the other, and the published findings came from the second question.

Where is HiltLock based, and who can you work with?

Operated from India by CalmSparks Tech Pvt. Ltd. Assessments run remotely against systems wherever they are hosted, and pricing is shown in rupees and dollars for that reason. We also work through consultancies and systems integrators who hold the client relationship themselves.

Will you tell us if we do not need you?

Yes, and it happens. A single-surface assistant with no tool access and no tenant boundary usually does not warrant a human red team. Saying so on the first call is worth more to us than one small engagement.

Judge us on the findings, not this page

Everything above is a claim. The evidence page is not.