What We Do
Five Lines, One Argument
The model is not the attack surface - everything it has permission to touch is. Every line here exists to find out what your systems can do, measure what they get wrong, or make sure somebody can prove it afterwards.
01
Security & Red-Teaming
Adversarial testing of what your AI systems can actually reach - prompt injection, tool-call abuse, data exfiltration - scoped by what an agent is able to do rather than by what it was told to do.
Includes
- Adversarial testing - prompt injection, jailbreaks, tool-call abuse
- Blast-radius review of every tool and API an agent can call
- Data-exfiltration paths - what enters context and what leaves
- Credential scoping and the identity an agent acts under
- Untrusted-content boundaries - retrieval sources, uploads, fetched pages
The model is not the attack surface. The attack surface is everything the model has been given permission to touch. An assistant that can only draft text is a writing tool with an embarrassing failure mode. An agent that can call your CRM, your payments API and your file store is a process running with credentials, and it takes instructions from whatever text happens to arrive in its context.
02
Accuracy & Evaluation
Measuring what your system actually gets wrong, on your own data, behind a regression gate that fails the build - rather than a demo that happened to work on the day it was shown.
Includes
- Eval sets built from your own data and your real edge cases
- Grounding checks - does the answer follow from the retrieved source
- Regression gates in CI, so a prompt change cannot ship silently
- Human review workflow for the cases automation cannot score
- An error taxonomy - what kind of wrong, how often, and where it costs
Most AI features ship on impressions. Somebody tried twenty prompts, the answers read well, and it went live. That is not a measurement. Twenty hand-picked prompts tell you the system can work; they tell you nothing about how often it does not, on which inputs, or whether last week's prompt edit made it quietly worse.
03
Governance & Policy
Written rules for what your AI systems may do, turned into controls that actually run - approval gates, scoped permissions, retention - plus the inventory and records an auditor or a customer will ask you for.
Includes
- AI system inventory - what exists, what data it touches, who owns it
- Usage policy written to be enforceable rather than aspirational
- Human-in-the-loop approval gates on irreversible actions
- Role and permission scoping for agent identities
- Evidence mapped to ISO/IEC 42001 and EU AI Act obligations
Policy documents fail in a predictable way. They describe intent, nothing enforces them, and the gap between the two only becomes visible during an incident or a security review. A rule that lives in a PDF is a preference.
04
Monitoring & Incident Response
Tracing every model call and tool call so you can say what a system did, on what input, and on whose authority - built before the incident, because it cannot be added afterwards.
Includes
- Tracing across model calls, retrieval and tool calls
- A reconstructable trail - input, context, version, authority, outcome
- Alerting on drift, refusals, latency, cost and anomalous tool use
- Incident runbooks and post-incident reconstruction
- OpenTelemetry-based, into the observability stack you already have
Most AI deployments cannot answer the first question anyone asks. A customer says your assistant told them something wrong on Tuesday. What was the input, which prompt version was live, which documents were retrieved, what did it go on to call, and who authorised that? In a lot of systems the honest answer is that nobody knows, because the only record is an application log that captured the response and discarded everything that produced it.
05
Efficiency & Cost Control
Cost and latency measured per outcome rather than per token - model routing, caching, context discipline - and finding the calls that spend money without changing the answer.
Includes
- Cost and latency attributed per outcome, not per token
- Model routing - the smallest model that still passes the evals
- Prompt-prefix caching and context trimming
- Retry, timeout and fallback policy that does not silently double spend
- Every efficiency change validated against the eval set first
AI spend behaves unlike most line items. It scales with usage, it arrives monthly with no unit economics attached, and very few teams can say what one completed task costs them. The invoice is a single number; the figure you actually need is cost per resolved ticket, per document processed, per outcome that mattered to somebody.
How It Runs
Evidence Before Recommendations
Nothing on this page gets scoped before the audit. Reading the systems first is the only way to price this work honestly - and the roadmap it produces is yours whether or not the rest follows.
01
Audit
A read-only diagnostic of what you run today: every AI system in use, what each can reach, where untrusted input can get to a privileged action, what data leaves, and whether anything is measured or recorded. It ends in a written roadmap ordered by risk.
02
Instrument
We build the things that have to exist before anything else can be judged - tracing through model and tool calls, an eval set from your own data, and the system inventory. Nothing here is a product of ours; it is built in your stack and stays there.
03
Govern & Operate
Controls go live - scoped permissions, approval gates on irreversible actions, alerting, regression gates in CI. Every week you get what changed, what the numbers did, and what we would do next. Controls that are not delivering value are removed.
04
Re-audit
Each cycle ends back at the diagnostic rather than at onboarding. Systems change, models get swapped, someone adds a tool - so what the last cycle established becomes the baseline for the next one, and the picture compounds instead of going stale.
Start With the Audit
A read-only review of every AI system you have in production - what each one can reach, what data leaves, and whether anything is measured or recorded. No fee, under mutual NDA, and the roadmap is yours either way.