Skip to content

Insights / 7 min Read

What Is AI Assurance? A Working Definition, and What It Is Not

AI assurance is evidence that an AI system does what it should, cannot do what it must not, and can be shown to have done either. What that means in practice, and how to tell if you have it.

Published 2026-09-19

AI assurance is the practice of producing evidence that an AI system does what it is meant to do, cannot do what it must not do, and can be shown afterwards to have done either - evidence good enough for the people who have to trust the system: the team that runs it, the customer that depends on it, and the auditor or regulator who eventually asks.

That is a working definition rather than a standards-body one, and it is deliberately built out of three questions, because in practice those three are the whole job. Everything sold under the name - red-teaming, evaluation, governance, monitoring - is a way of answering one of them.

Where the word comes from

Assurance is borrowed from audit and from safety engineering, where it has a precise meaning: independent evidence that a thing meets a standard, as distinct from the thing being claimed to. An accountant's assurance over a set of accounts is not a promise that the business is good; it is a statement that the numbers were checked and by what method. A safety case for a bridge is not a guarantee it will never fall; it is the documented argument, with evidence, for why it is expected to stand.

Applied to AI the word carries the same discipline, and the same limit. Assurance does not make a system correct or safe. It establishes, with evidence rather than confidence, what the system does and does not do - which is the thing most organisations running AI today cannot actually state.

The three questions

1. Does it do what it should?

This is accuracy, and it is answered by measurement rather than by demonstration. A set of real inputs with known-correct outputs, drawn from your own domain, run against the system on every change, with a number somebody is accountable for. For anything that answers from documents it also means checking grounding - whether each claim follows from the retrieved text - separately from whether the answer reads well. Twenty prompts that looked right in a demo do not answer this question. They answer a different one, which is whether the system can work at all.

2. Can it do what it must not?

This is security and permissions, and it is the question that changed shape when models started calling tools. As long as an AI system only produced text for a person to read, the person was the control: a wrong or malicious answer had to get past a human before it did anything. An agent that can call an API, send an email, query a database or move money has no such person in the way, and it takes instructions from whatever text arrives in its context - a ticket, a web page, a document a customer uploaded. Assurance here means knowing what each system is technically able to reach, testing what an attacker who controls some of its input could make it do, and bounding that with permissions, approval gates and architecture rather than with prompt wording.

3. Can you show what it did?

This is governance and records. For any run the system performed: what was the input, which prompt and model versions were live, what was retrieved, what tools were called and with what arguments, on whose authority, and what came out. If that cannot be reconstructed, the system cannot be debugged after an incident, defended to a customer, or audited against any framework - and it has to be instrumented before the event, because nothing about an AI system is rerunnable afterwards. Alongside the trace sits the inventory: which AI systems exist at all, what data each touches, who owns it, and what it is permitted to do.

What AI assurance is not

  • It is not AI safety research. That field works on the behaviour of frontier models themselves. Assurance takes the model as given and works on the system built around it.
  • It is not AI ethics. Ethics is about which uses are acceptable. Assurance is about whether a system does what was decided, and whether that can be shown.
  • It is not certification. A certificate is a third party's statement that a standard was met. Assurance is the evidence that statement would rest on. A firm that prepares you for an audit should not also be the one certifying you.
  • It is not a product you install. Tools help - tracing backends, eval harnesses, policy engines - but assurance is the evidence those tools produce about your specific system, and no tool produces it by being present.
  • It is not model selection. Which model you use changes the answers less than what the model is permitted to do, what it is grounded on, and whether anyone is measuring.

Assurance and governance are not the same thing

The two words get used interchangeably and they should not be. Governance is the rules and the accountability: what AI systems may do, who decides, who owns each one, what is recorded. Assurance is the evidence that those rules actually hold - that the permission scope written in the policy is the one the agent runs with, that the accuracy claimed in the review is the accuracy measured on real inputs, that the approval gate fires before the irreversible action and not after.

Governance without assurance is a policy document. It describes an intended state, nothing verifies it, and the gap is discovered during an incident. Assurance without governance is a pile of measurements nobody is accountable for acting on. An organisation needs both, and the honest order is to measure first, because a policy written before anyone knows what the systems can actually do tends to govern the systems the authors imagined rather than the ones that exist.

Why it has become urgent

Three things arrived at once. Agents, which removed the human from between the model and the action. Regulation - ISO/IEC 42001 as a certifiable management standard, the EU AI Act with obligations phasing in through 2027, sector rules, and customer security questionnaires that now ask about AI specifically - all of which want records rather than reassurance. And spend, which has reached the size where somebody in finance asks what one completed task costs and nobody can answer.

The common thread is that all three are questions of evidence. Not whether the model is good, but whether you can show what your system does.

What it looks like in practice

In our work it is five lines, each answering one of the questions above or making the answer cheaper: security and red-teaming, accuracy and evaluation, governance and policy, monitoring and incident response, and efficiency and cost control. An engagement starts with a read-only audit of what is already running, because scoping the work honestly requires reading the systems first, and the audit itself tends to answer the question of which line matters most - which is frequently not the one the team came in asking about.

How to tell whether you already have it

Three tests, each answerable in an afternoon.

  • Pick one real interaction from last week and reconstruct it completely: input, versions, retrieved documents, tool calls, authorising identity, output. If you get stuck at the third item, the system cannot be audited.
  • State the current error rate of your most important AI feature, on your own data, and name the person who watches it. If the answer is a demo or an impression, accuracy is not being measured.
  • List every tool, API and credential each agent can reach, and mark which of those actions are irreversible. If nobody can produce the list, nobody knows the blast radius.

An organisation that can do all three has assurance, whatever it calls it. One that cannot has hope, and the point of the word is to stop confusing the two.

General information, not legal or professional advice about your organisation. Regulatory obligations depend on specific facts; take advice from qualified counsel before acting. See the Terms of Use.

Questions This Raises

What is AI assurance, in one sentence?

AI assurance is evidence - measured, recorded and reconstructable - that an AI system does what it is meant to, cannot do what it must not, and can be shown afterwards to have done either.

Is AI assurance the same as AI governance?

No. Governance is the rules and accountability - what systems may do, who owns them, what is recorded. Assurance is the evidence that those rules actually hold in the running system. Governance without assurance is a policy document; assurance without governance is measurement nobody acts on.

Who is responsible for AI assurance in a company?

In practice the team that ships the AI feature owns the measurement and the instrumentation, security owns the permission and red-teaming questions, and whoever is accountable for risk or compliance owns the inventory and the records. Where it fails is when each assumes another has it. The first useful step is usually naming one person who can answer all three questions for each system.

Is AI assurance required by law?

Not under that name, but the evidence it produces is what several frameworks ask for. ISO/IEC 42001 requires a managed inventory, risk treatment and records; the EU AI Act imposes logging, documentation and human-oversight obligations on regulated systems, phasing in through 2027; and customer security reviews increasingly ask for the same things. Assurance is how those records come to exist before they are demanded.

Want This Applied To Your Systems?

The initial audit applies this to the systems you run and produces a written roadmap ordered by risk.

Book a Call