Skip to content

Sample Audit Report

What The Audit Hands Over

An illustrative report for a fictional company, in the format every no-fee audit is delivered in. The company, its systems and its findings are invented; the structure is exactly what you would receive.

Illustrative only. "Example Lending" is a fictional company. No client's systems, data or findings appear on this page, and no client's report is ever published.

01 · Summary

Summary

Example Lending runs three AI systems in production. One of them - the customer support agent - can move money, reads text written by customers, and has nothing between the model's decision and the payment. That combination is the single most important finding in this report, and it can be contained this week without rebuilding anything.

  • 1 critical, 2 high, 2 medium, 1 low, ordered by what we could demonstrate, not by label.
  • Accuracy is not measured on any system, and no past interaction could be reconstructed.
  • Three fixes are recommended for this week; all three are small changes to existing code and settings, not new systems.

02 · Scope

Scope And Access

  • Read-only access to two source repositories, the model provider console, application logs and the IAM roles the agents run as.
  • The one adversarial test (F2) ran on staging, under written authorisation that named the system, the window and the technique.
  • No production data was copied out of the client's environment. All access was revoked at the end of the audit and revocation was confirmed in writing.

03 · Inventory

AI Systems Found

SystemWhat it doesWhat it can reachOwner
Support agentAnswers tickets, issues refundsPayments API (write), CRM (read/write)Head of Support
Policy assistantAnswers staff questions from internal policy documentsDocument index (read)HR Operations
Developer copilotCode suggestionsSource repositories (read)None - found in audit

04 · Findings

Findings

Each finding states what we observed, what it lets happen, how to fix it, and the framework reference your risk team will want. Ordered by demonstrated impact.

F1 · Critical · Blast radius

Support agent can issue refunds with no approval step

Evidence
The support agent's service account holds a payments API key scoped to refunds:write with no amount ceiling. The tool definition passes the model-chosen amount straight to the API. No human approval, rate limit or per-ticket cap sits between the model's decision and the refund.
Impact
Any path that steers the model (see F2) converts directly into money leaving the business. Refunds are irreversible once settled.
Remediation
Cap the tool at the order value it is refunding, route anything above a threshold to a human queue, and replace the shared key with a scoped key that cannot exceed the cap even if the tool code is bypassed.
Maps To
OWASP LLM06 Excessive Agency

F2 · High · Untrusted input paths

Customer-written ticket text reaches the refund tool unfiltered

Evidence
Ticket bodies and attachments are inserted into the agent's context verbatim. In a test ticket on the staging environment, an instruction embedded in an attached PDF caused the agent to propose a refund the ticket did not request.
Impact
An outsider controls text that sits next to a privileged tool. Combined with F1, this is the highest-priority path in the estate.
Remediation
Fixing F1 bounds the damage this week. Then separate reading from acting: the model that reads untrusted text proposes, and a policy check on structured fields - not on the model's reasoning - decides.
Maps To
OWASP LLM01 Prompt Injection · MITRE ATLAS AML.T0051

F3 · High · Data exposure and retention

Customer PII sent to a model provider with training use not confirmed off

Evidence
Full ticket threads, including phone numbers and partial account numbers, are sent to the provider. The account is on a plan tier whose data-use terms were not reviewed, and nobody could confirm where the opt-out setting stood.
Impact
Personal data may be retained or used outside the terms the company has given its customers.
Remediation
Confirm and record the provider's data-use setting, redact identifiers before the model call, and add the provider to the processor register.
Maps To
OWASP LLM02 Sensitive Information Disclosure · DPDP Act, 2023

F4 · Medium · Whether actions are recorded

No record of which prompt version produced a given answer

Evidence
We selected one escalated conversation from the previous week and could not reconstruct the prompt version, the retrieved documents or the tool calls. Application logs record the final reply only.
Impact
After an incident the company could not show what happened or why, to a customer or to itself.
Remediation
Emit one trace per interaction with prompt version, retrieved document IDs, tool calls and the authorising identity. Retention to match the complaint window.
Maps To
NIST AI RMF - Manage · ISO/IEC 42001

F5 · Medium · Whether accuracy is measured at all

Answer accuracy is not measured

Evidence
No evaluation set exists. The policy assistant was approved on a demo. Prompt changes ship without any regression check.
Impact
The error rate is unknown, so there is no way to tell whether a change made it better or worse.
Remediation
Build a first eval set of 100-150 real questions with approved answers, and fail the build when grounded accuracy drops.
Maps To
OWASP LLM09 Misinformation · NIST AI RMF - Measure

F6 · Low · What is actually running

A developer copilot was not on the AI register

Evidence
A code assistant is enabled organisation-wide through an existing SaaS licence. It was not on the register and had no named owner.
Impact
Low on its own, but it means the register cannot be relied on for the rest of the estate.
Remediation
Add it with an owner, record its data settings, and make the register the gate for switching on any new AI feature.
Maps To
ISO/IEC 42001 - AI system inventory

05 · This Week

Fixes For This Week

  • Cap and gate refunds (F1) - a per-ticket ceiling in the tool, and human approval above a threshold.
  • Replace the payments key (F1) with one scoped to the cap, so the limit holds even if the tool code is bypassed.
  • Confirm the provider's data-use setting (F3) and record it in the processor register.

These three contain the critical path. They do not fix F2 - they make it cheap.

06 · Roadmap

Roadmap

Next 30 days

  • Separate reading from acting on the support agent (F2).
  • Redact identifiers before the model call (F3).
  • One trace per interaction, with prompt version and tool calls (F4).

Next 90 days

  • A first eval set and a regression gate on prompt changes (F5).
  • The AI register as the gate for switching on new AI features (F6).

Each item says why it is where it is, so the order can be argued with. If you disagree with it, the report is still yours to use without us.

07 · Limitations

Limitations

This report describes the systems as they were configured during the audit window. Testing shows whether specific attacks work; it does not prove that no others do. It is not a certification against any framework - see Independence And Limitations.

Get One For Your Systems

The same report, on your systems - read-only, at no cost, under mutual NDA, within 48 hours of access.

Book a Call