Service 01
Security & Red-Teaming
Adversarial testing of what your AI systems can actually reach - prompt injection, tool-call abuse, data exfiltration - scoped by what an agent is able to do rather than by what it was told to do.
The model is not the attack surface. The attack surface is everything the model has been given permission to touch. An assistant that can only draft text is a writing tool with an embarrassing failure mode. An agent that can call your CRM, your payments API and your file store is a process running with credentials, and it takes instructions from whatever text happens to arrive in its context.
Prompt injection is the concrete version of that problem. Untrusted content - a support ticket, a fetched web page, a PDF a customer uploaded, a row returned from a database - carries instructions, and a model has no reliable way to separate content it was asked to read from an instruction addressed to it. Defences written at the prompt layer lower the rate. They do not close it, and a control that holds most of the time is not a control.
So we test where the damage would land rather than where the prompt is written: what each set of agent credentials can reach, which tool calls are irreversible, what data crosses the boundary and where it goes afterwards. The output is an ordered list of findings, each with the blast radius we could actually demonstrate - not a score out of ten.
Questions About This Work
What is AI red teaming?
Adversarial testing of an AI system to find out what it can be made to do that it should not, by an attacker who controls some of the text it reads. For an LLM application that covers prompt injection, jailbreaks, tool-call abuse and data exfiltration. The useful output is not a score out of ten; it is the list of what could actually be reached, ordered by the damage it would do.
How is LLM red teaming different from a normal penetration test?
A conventional penetration test probes code and infrastructure that behave deterministically. An LLM application also takes instructions from its inputs, so the attack surface includes every document, page, ticket or record the model reads, and the findings are about what the model was permitted to do with them rather than about a patchable bug. The two overlap and both are needed; neither substitutes for the other.
Do you test against the OWASP Top 10 for LLM Applications?
Yes. It is a sensible checklist and we map findings to it so they are legible to a security team. But the work is scoped by your agents' actual permissions rather than by the list, because the list tells you which classes of failure exist and your blast radius tells you which of them can hurt you.
What do we get at the end of a red-teaming engagement?
An ordered list of findings, each with the path we followed, the blast radius we could demonstrate, and a fix stated in terms of permissions or architecture rather than prompt wording. Plus the untrusted-input map and the credential inventory built along the way, which are yours to keep.
Reading On This
Start With the Audit
No engagement is scoped before a read-only review of the systems you actually run. The roadmap it produces is yours either way.