Skip to content

Insights / 7 min Read

What an AI Audit Trail Has to Contain to Be Worth Anything

Most AI logging records the answer and throws away everything that produced it. Here is what a trail needs to contain to survive an incident or an audit.

Published 2026-09-10

Here is a test worth running this week. Pick one real interaction your AI system had a few days ago - ideally one somebody complained about. Now reconstruct it completely: the exact input, the prompt version live at that moment, the model and its version, the documents retrieved, the tools called and with what arguments, the identity that authorised those calls, and the final output.

Most teams cannot get past the third item. The application logged a request and a response, which felt like logging at the time, and everything that explains the response is gone.

Why the usual logging is not enough

Conventional application logging assumes deterministic code. If you know the inputs and the version of the code, you can rerun it and get the same result, so the log only has to record enough to find the entry point. Everything else is recoverable by reading the source.

None of that holds here. The same input to the same model can produce different output. The behaviour depends on a prompt that is data rather than code, and that may have been edited by someone who never opened a pull request. It depends on documents retrieved from an index whose contents change hourly. It depends on a model version the provider can update underneath you. Nothing is rerunnable, so nothing is recoverable after the fact - if it was not captured when it happened, it does not exist.

The six things a trail has to carry

1. The input, before anything touched it

The raw request as it arrived, separately from the assembled prompt. When those two are conflated you cannot tell whether a bad answer came from a bad question or from your own preprocessing mangling a good one.

2. Every version in play

Prompt version, model identifier and version, retrieval index version, tool schema version, application release. This is the field most often missing and the one that costs the most, because without it "did this get worse after Tuesday" is unanswerable. Version your prompts like code, in a repository, with a history - a prompt edited in a console with no record is an untracked deployment to production.

3. The retrieved context

Which documents were retrieved, their identifiers and versions, their scores, and what was actually placed in the context window after truncation. Truncation matters more than people expect: a system that retrieved the correct document and then cut it off before the relevant paragraph looks identical, from the outside, to one that retrieved the wrong document entirely. Storing document ids and hashes rather than full text keeps this affordable.

4. Every tool call, with arguments and results

What was called, with which arguments, what came back, whether it succeeded, how long it took. For anything that changes state, this is the difference between an audit trail and an anecdote. This is also where the agent stops being a text generator and starts being a process that does things, which is exactly where the record needs to be strongest.

5. The authorising identity

On whose authority did each action happen - which user, which agent identity, which credential, and if a human approved it, who, when, and what were they shown at the moment they approved. "The system did it" is not an answer that survives either an incident review or an auditor, and an approval record that does not capture what the approver saw is close to worthless.

6. The link that ties it together

All of the above sharing one trace id, so a run is a single connected object rather than six sources someone correlates by timestamp in the middle of an incident. This is the part that decides whether an investigation takes ten minutes or two days.

Use OpenTelemetry, not a bespoke format

The temptation is to write a custom logger, because the requirements above look specific. Do not. Traces and spans already model exactly this shape - a run is a trace, each model call, retrieval and tool call is a span, and parent-child relationships capture the structure of an agent loop natively.

The practical arguments are stronger than the aesthetic one. Your AI traces land in the same system as the rest of your telemetry, so an investigation does not stop at the boundary of the AI feature. You get sampling, batching, redaction and context propagation without building them. Semantic conventions for model calls exist and are stabilising, so field names are not invented locally and re-invented differently by the next team. And you can change backends without re-instrumenting, which matters because most teams pick the wrong one first.

What to do about sensitive content

The obvious objection: prompts and outputs contain personal and confidential data, and a trail like this is a copy of all of it in a second system. That is a real problem and it has ordinary answers - redact structured identifiers at the point of capture, store hashes and ids rather than full document text, keep a short retention window on payloads while keeping metadata and versions for far longer, and put access controls on the trace store equal to those on the source data.

What is not a good answer is logging nothing, which is the default a surprising number of teams arrive at by never deciding. Metadata alone - versions, tool calls, identities, timings, document ids - is enormously useful, carries far less sensitive content than the payloads, and is the part that answers "what changed and when".

Where this pays off

Three places, and only the third is the one people build it for. Debugging, where you go from arguing about a screenshot to reading what happened. Incidents, where the question is what else was affected, and an un-instrumented system forces you to answer "we cannot rule anything out" - which is the answer that turns a small incident into a large disclosure. And audit, where an ISO/IEC 42001 or EU AI Act conversation, or simply a customer's security questionnaire, asks for records of what your systems did and who authorised it.

The uncomfortable property of all three is that the work has to be done before it is needed, and it is invisible until then. That is why it gets deferred, and why the teams who did it early are so noticeably calmer during the week when it matters.

General information, not legal or professional advice about your organisation. Regulatory obligations depend on specific facts; take advice from qualified counsel before acting. See the Terms of Use.

Questions This Raises

Why is normal application logging not enough for AI systems?

Conventional logging assumes you can rerun deterministic code from the same inputs. AI systems are not rerunnable - output varies, prompts are data that may change outside version control, retrieval indexes change hourly, and providers update model versions underneath you. Anything not captured at the moment it happened cannot be recovered afterwards by reading the source.

What should an AI audit trail record?

Six things: the raw input before preprocessing; every version in play (prompt, model, index, tool schema, release); the retrieved context including what survived truncation; every tool call with arguments and results; the authorising identity, including who approved an action and what they were shown; and a shared trace id linking all of it into one reconstructable run.

Should we use OpenTelemetry for LLM observability?

In most cases yes. Traces and spans already model an agent run natively - a run is a trace, each model call, retrieval and tool call is a span. You inherit sampling, batching, redaction and context propagation, your AI telemetry sits alongside the rest of your system, semantic conventions for model calls exist and are stabilising, and you can change backend without re-instrumenting.

How do we log AI interactions without storing sensitive data?

Redact structured identifiers at capture time, store document ids and hashes rather than full text, keep payloads on a short retention window while keeping metadata and versions much longer, and apply the same access controls to the trace store as to the source data. Metadata alone answers most questions about what changed and when, and carries far less exposure than the payloads.

Want This Applied To Your Systems?

The initial audit applies this to the systems you run and produces a written roadmap ordered by risk.

Book a Call