<?xml version="1.0" encoding="UTF-8"?>
<!-- Generated by scripts/prerender.mjs from ARTICLES. Do not edit by hand. -->
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:dc="http://purl.org/dc/elements/1.1/">
<channel>
  <title>Vetro Insights</title>
  <link>https://vetro.co.in/insights</link>
  <description>Long-form, specific writing on prompt injection, AI audit trails and measuring accuracy before you ship - the questions clients ask before they hire anyone.</description>
  <language>en</language>
  <atom:link href="https://vetro.co.in/feed.xml" rel="self" type="application/rss+xml" />
  <lastBuildDate>Sat, 19 Sep 2026 03:30:00 GMT</lastBuildDate>
  <item>
    <title>What Is AI Assurance? A Working Definition, and What It Is Not</title>
    <link>https://vetro.co.in/insights/what-is-ai-assurance</link>
    <guid isPermaLink="true">https://vetro.co.in/insights/what-is-ai-assurance</guid>
    <pubDate>Sat, 19 Sep 2026 03:30:00 GMT</pubDate>
    <dc:creator>Prithviraj Chawla</dc:creator>
    <description>AI assurance is evidence that an AI system does what it should, cannot do what it must not, and can be shown to have done either. What that means in practice, and how to tell if you have it.</description>
    <content:encoded><![CDATA[<p>AI assurance is the practice of producing evidence that an AI system does what it is meant to do, cannot do what it must not do, and can be shown afterwards to have done either - evidence good enough for the people who have to trust the system: the team that runs it, the customer that depends on it, and the auditor or regulator who eventually asks.</p>
<p>That is a working definition rather than a standards-body one, and it is deliberately built out of three questions, because in practice those three are the whole job. Everything sold under the name - red-teaming, evaluation, governance, monitoring - is a way of answering one of them.</p>
<h2>Where the word comes from</h2>
<p>Assurance is borrowed from audit and from safety engineering, where it has a precise meaning: independent evidence that a thing meets a standard, as distinct from the thing being claimed to. An accountant's assurance over a set of accounts is not a promise that the business is good; it is a statement that the numbers were checked and by what method. A safety case for a bridge is not a guarantee it will never fall; it is the documented argument, with evidence, for why it is expected to stand.</p>
<p>Applied to AI the word carries the same discipline, and the same limit. Assurance does not make a system correct or safe. It establishes, with evidence rather than confidence, what the system does and does not do - which is the thing most organisations running AI today cannot actually state.</p>
<h2>The three questions</h2>
<h3>1. Does it do what it should?</h3>
<p>This is accuracy, and it is answered by measurement rather than by demonstration. A set of real inputs with known-correct outputs, drawn from your own domain, run against the system on every change, with a number somebody is accountable for. For anything that answers from documents it also means checking grounding - whether each claim follows from the retrieved text - separately from whether the answer reads well. Twenty prompts that looked right in a demo do not answer this question. They answer a different one, which is whether the system can work at all.</p>
<h3>2. Can it do what it must not?</h3>
<p>This is security and permissions, and it is the question that changed shape when models started calling tools. As long as an AI system only produced text for a person to read, the person was the control: a wrong or malicious answer had to get past a human before it did anything. An agent that can call an API, send an email, query a database or move money has no such person in the way, and it takes instructions from whatever text arrives in its context - a ticket, a web page, a document a customer uploaded. Assurance here means knowing what each system is technically able to reach, testing what an attacker who controls some of its input could make it do, and bounding that with permissions, approval gates and architecture rather than with prompt wording.</p>
<h3>3. Can you show what it did?</h3>
<p>This is governance and records. For any run the system performed: what was the input, which prompt and model versions were live, what was retrieved, what tools were called and with what arguments, on whose authority, and what came out. If that cannot be reconstructed, the system cannot be debugged after an incident, defended to a customer, or audited against any framework - and it has to be instrumented before the event, because nothing about an AI system is rerunnable afterwards. Alongside the trace sits the inventory: which AI systems exist at all, what data each touches, who owns it, and what it is permitted to do.</p>
<h2>What AI assurance is not</h2>
<ul><li>It is not AI safety research. That field works on the behaviour of frontier models themselves. Assurance takes the model as given and works on the system built around it.</li><li>It is not AI ethics. Ethics is about which uses are acceptable. Assurance is about whether a system does what was decided, and whether that can be shown.</li><li>It is not certification. A certificate is a third party's statement that a standard was met. Assurance is the evidence that statement would rest on. A firm that prepares you for an audit should not also be the one certifying you.</li><li>It is not a product you install. Tools help - tracing backends, eval harnesses, policy engines - but assurance is the evidence those tools produce about your specific system, and no tool produces it by being present.</li><li>It is not model selection. Which model you use changes the answers less than what the model is permitted to do, what it is grounded on, and whether anyone is measuring.</li></ul>
<h2>Assurance and governance are not the same thing</h2>
<p>The two words get used interchangeably and they should not be. Governance is the rules and the accountability: what AI systems may do, who decides, who owns each one, what is recorded. Assurance is the evidence that those rules actually hold - that the permission scope written in the policy is the one the agent runs with, that the accuracy claimed in the review is the accuracy measured on real inputs, that the approval gate fires before the irreversible action and not after.</p>
<p>Governance without assurance is a policy document. It describes an intended state, nothing verifies it, and the gap is discovered during an incident. Assurance without governance is a pile of measurements nobody is accountable for acting on. An organisation needs both, and the honest order is to measure first, because a policy written before anyone knows what the systems can actually do tends to govern the systems the authors imagined rather than the ones that exist.</p>
<h2>Why it has become urgent</h2>
<p>Three things arrived at once. Agents, which removed the human from between the model and the action. Regulation - ISO/IEC 42001 as a certifiable management standard, the EU AI Act with obligations phasing in through 2027, sector rules, and customer security questionnaires that now ask about AI specifically - all of which want records rather than reassurance. And spend, which has reached the size where somebody in finance asks what one completed task costs and nobody can answer.</p>
<p>The common thread is that all three are questions of evidence. Not whether the model is good, but whether you can show what your system does.</p>
<h2>What it looks like in practice</h2>
<p>In our work it is five lines, each answering one of the questions above or making the answer cheaper: security and red-teaming, accuracy and evaluation, governance and policy, monitoring and incident response, and efficiency and cost control. An engagement starts with a read-only audit of what is already running, because scoping the work honestly requires reading the systems first, and the audit itself tends to answer the question of which line matters most - which is frequently not the one the team came in asking about.</p>
<h2>How to tell whether you already have it</h2>
<p>Three tests, each answerable in an afternoon.</p>
<ul><li>Pick one real interaction from last week and reconstruct it completely: input, versions, retrieved documents, tool calls, authorising identity, output. If you get stuck at the third item, the system cannot be audited.</li><li>State the current error rate of your most important AI feature, on your own data, and name the person who watches it. If the answer is a demo or an impression, accuracy is not being measured.</li><li>List every tool, API and credential each agent can reach, and mark which of those actions are irreversible. If nobody can produce the list, nobody knows the blast radius.</li></ul>
<p>An organisation that can do all three has assurance, whatever it calls it. One that cannot has hope, and the point of the word is to stop confusing the two.</p>]]></content:encoded>
  </item>
  <item>
    <title>Measuring an AI Feature: What to Check Before You Ship It</title>
    <link>https://vetro.co.in/insights/ai-evaluation-before-you-ship</link>
    <guid isPermaLink="true">https://vetro.co.in/insights/ai-evaluation-before-you-ship</guid>
    <pubDate>Thu, 10 Sep 2026 03:30:00 GMT</pubDate>
    <dc:creator>Prithviraj Chawla</dc:creator>
    <description>Twenty prompts that looked right is not a measurement. What an eval set needs to contain, what to measure, and how to stop regressions reaching users.</description>
    <content:encoded><![CDATA[<p>Almost every AI feature ships the same way. Someone builds it, tries a couple of dozen prompts, the answers read well, a demo goes smoothly, and it goes live. Then it is in production and nobody can say how often it is wrong, on what kind of input, or whether this week's prompt edit improved anything.</p>
<p>The gap is not rigour for its own sake. It is that the twenty prompts were written by the person who built the thing, which means they are the inputs it was implicitly designed to handle. They demonstrate that the system can work. They say nothing about how often it does.</p>
<h2>The eval set is the whole thing</h2>
<p>An eval set is a collection of real inputs with known-correct outputs. Everything else in evaluation is technique; this is the part that decides whether any of it means anything, and it is the part teams try hardest to skip because it is manual and dull.</p>
<h3>Draw from real traffic, not imagination</h3>
<p>The single highest-value move is to build the set from inputs users actually sent, once you have any. Real inputs are misspelt, under-specified, pasted from elsewhere, several questions at once, and occasionally in a language nobody planned for. Synthetic inputs written by the team are uniformly better-formed than reality, which is precisely why a system tuned on them is disappointing in production.</p>
<h3>Weight it towards the awkward cases</h3>
<p>A set that mirrors production's distribution will be mostly easy cases, which wastes most of the measurement on things that were never going to fail. Deliberately over-sample the edges: questions with no answer in the corpus, where the honest response is to say so; ambiguous questions with two defensible readings; inputs near a policy boundary; long inputs that stress the context window; adversarial ones. The middle of the distribution is not where you get hurt.</p>
<h3>A hundred good cases beats a thousand careless ones</h3>
<p>Every case needs an output somebody knowledgeable has actually confirmed is right. That is slow, and it is also the source of all the value - an eval set with unreliable labels produces confident numbers that mean nothing, which is worse than having no numbers at all. Start at a hundred, curated carefully, and grow it from real failures as they surface.</p>
<h2>Measure grounding separately from everything else</h2>
<p>For any system that answers from documents, the most useful single measurement is whether each claim in the answer actually follows from the retrieved text. Not whether it is true in general, and not whether it reads well - whether it is supported by what was in front of the model.</p>
<p>Keeping this separate matters because it splits one vague complaint - "it hallucinates" - into two different bugs with different fixes. If the right document was retrieved and the answer still is not supported by it, that is a generation problem: prompt, model, context assembly. If the right document was never retrieved, no amount of prompt work fixes it, and every hour spent there is wasted. Measure retrieval and generation separately or you will spend your effort in the wrong half.</p>
<p>The other measurement worth having early: on questions the corpus genuinely cannot answer, does the system say it does not know? A system that never abstains has not been shown to be accurate, only to be confident.</p>
<h2>On using a model as the judge</h2>
<p>Scoring answers by hand does not scale, so the standard approach is to have a model grade them. It works well enough to be worth doing, with three caveats worth knowing before you trust the number.</p>
<ul><li>Judges have biases - towards longer answers, towards their own family's phrasing, towards the first option shown in a pairwise comparison. Randomise order and be careful comparing across model families.</li><li>A judge scoring a vague criterion like "helpfulness" produces a number with no defensible meaning. Narrow, checkable questions - "is every claim supported by the source text, yes or no" - are far more reliable.</li><li>A judge needs its own validation. Score a hundred cases by hand, compare against the judge, and measure the agreement. If you have not done this, you do not know what your eval numbers mean.</li></ul>
<p>And keep a human in the loop for the cases automation genuinely cannot score. There is always a residue, and pretending otherwise is how a metric drifts away from the thing it was supposed to represent.</p>
<h2>Wire it to the build, or it will rot</h2>
<p>An eval suite that has to be run manually gets run before big launches and forgotten in between, which is exactly backwards - the regressions that hurt come from small changes nobody thought needed checking.</p>
<p>So it runs automatically on any change to a prompt, a model version, a retrieval configuration, a tool schema or a routing rule, and it fails the build when a metric drops past a threshold agreed in advance. Agreed in advance is the operative part. A threshold negotiated after a red build is not a gate, it is a conversation, and it always resolves in favour of shipping.</p>
<p>One deliberate exception is worth building in: expected changes. When a prompt is meant to alter behaviour, the suite should make it easy to review the differences and accept them explicitly, rather than tempting anyone to switch the gate off.</p>
<h2>Then keep measuring in production</h2>
<p>An eval set is a fixed sample of a world that moves. Users ask new things, the corpus changes, providers update models underneath you. Offline evaluation catches regressions you caused; it does not catch drift you did not.</p>
<p>The production signals that pay for themselves are cheap: refusal and error rates, latency and cost per outcome, retrieval scores trending down, and thin implicit feedback like whether a user rephrased the same question immediately. Feed real failures back into the eval set as they surface. That loop - production failure becomes a permanent test case - is what stops the same bug shipping twice, and it is worth more than any single metric on the dashboard.</p>
<h2>The realistic minimum</h2>
<p>If you take one thing: a hundred real cases with verified answers, a grounding check, a threshold agreed before the build turns red, and a habit of adding every production failure back to the set. That is a weekend of work and about a week of labelling, and it is the difference between knowing what your system does and believing what your demo did.</p>]]></content:encoded>
  </item>
  <item>
    <title>What an AI Audit Trail Has to Contain to Be Worth Anything</title>
    <link>https://vetro.co.in/insights/ai-audit-trail</link>
    <guid isPermaLink="true">https://vetro.co.in/insights/ai-audit-trail</guid>
    <pubDate>Thu, 10 Sep 2026 03:30:00 GMT</pubDate>
    <dc:creator>Prithviraj Chawla</dc:creator>
    <description>Most AI logging records the answer and throws away everything that produced it. Here is what a trail needs to contain to survive an incident or an audit.</description>
    <content:encoded><![CDATA[<p>Here is a test worth running this week. Pick one real interaction your AI system had a few days ago - ideally one somebody complained about. Now reconstruct it completely: the exact input, the prompt version live at that moment, the model and its version, the documents retrieved, the tools called and with what arguments, the identity that authorised those calls, and the final output.</p>
<p>Most teams cannot get past the third item. The application logged a request and a response, which felt like logging at the time, and everything that explains the response is gone.</p>
<h2>Why the usual logging is not enough</h2>
<p>Conventional application logging assumes deterministic code. If you know the inputs and the version of the code, you can rerun it and get the same result, so the log only has to record enough to find the entry point. Everything else is recoverable by reading the source.</p>
<p>None of that holds here. The same input to the same model can produce different output. The behaviour depends on a prompt that is data rather than code, and that may have been edited by someone who never opened a pull request. It depends on documents retrieved from an index whose contents change hourly. It depends on a model version the provider can update underneath you. Nothing is rerunnable, so nothing is recoverable after the fact - if it was not captured when it happened, it does not exist.</p>
<h2>The six things a trail has to carry</h2>
<h3>1. The input, before anything touched it</h3>
<p>The raw request as it arrived, separately from the assembled prompt. When those two are conflated you cannot tell whether a bad answer came from a bad question or from your own preprocessing mangling a good one.</p>
<h3>2. Every version in play</h3>
<p>Prompt version, model identifier and version, retrieval index version, tool schema version, application release. This is the field most often missing and the one that costs the most, because without it "did this get worse after Tuesday" is unanswerable. Version your prompts like code, in a repository, with a history - a prompt edited in a console with no record is an untracked deployment to production.</p>
<h3>3. The retrieved context</h3>
<p>Which documents were retrieved, their identifiers and versions, their scores, and what was actually placed in the context window after truncation. Truncation matters more than people expect: a system that retrieved the correct document and then cut it off before the relevant paragraph looks identical, from the outside, to one that retrieved the wrong document entirely. Storing document ids and hashes rather than full text keeps this affordable.</p>
<h3>4. Every tool call, with arguments and results</h3>
<p>What was called, with which arguments, what came back, whether it succeeded, how long it took. For anything that changes state, this is the difference between an audit trail and an anecdote. This is also where the agent stops being a text generator and starts being a process that does things, which is exactly where the record needs to be strongest.</p>
<h3>5. The authorising identity</h3>
<p>On whose authority did each action happen - which user, which agent identity, which credential, and if a human approved it, who, when, and what were they shown at the moment they approved. "The system did it" is not an answer that survives either an incident review or an auditor, and an approval record that does not capture what the approver saw is close to worthless.</p>
<h3>6. The link that ties it together</h3>
<p>All of the above sharing one trace id, so a run is a single connected object rather than six sources someone correlates by timestamp in the middle of an incident. This is the part that decides whether an investigation takes ten minutes or two days.</p>
<h2>Use OpenTelemetry, not a bespoke format</h2>
<p>The temptation is to write a custom logger, because the requirements above look specific. Do not. Traces and spans already model exactly this shape - a run is a trace, each model call, retrieval and tool call is a span, and parent-child relationships capture the structure of an agent loop natively.</p>
<p>The practical arguments are stronger than the aesthetic one. Your AI traces land in the same system as the rest of your telemetry, so an investigation does not stop at the boundary of the AI feature. You get sampling, batching, redaction and context propagation without building them. Semantic conventions for model calls exist and are stabilising, so field names are not invented locally and re-invented differently by the next team. And you can change backends without re-instrumenting, which matters because most teams pick the wrong one first.</p>
<h2>What to do about sensitive content</h2>
<p>The obvious objection: prompts and outputs contain personal and confidential data, and a trail like this is a copy of all of it in a second system. That is a real problem and it has ordinary answers - redact structured identifiers at the point of capture, store hashes and ids rather than full document text, keep a short retention window on payloads while keeping metadata and versions for far longer, and put access controls on the trace store equal to those on the source data.</p>
<p>What is not a good answer is logging nothing, which is the default a surprising number of teams arrive at by never deciding. Metadata alone - versions, tool calls, identities, timings, document ids - is enormously useful, carries far less sensitive content than the payloads, and is the part that answers "what changed and when".</p>
<h2>Where this pays off</h2>
<p>Three places, and only the third is the one people build it for. Debugging, where you go from arguing about a screenshot to reading what happened. Incidents, where the question is what else was affected, and an un-instrumented system forces you to answer "we cannot rule anything out" - which is the answer that turns a small incident into a large disclosure. And audit, where an ISO/IEC 42001 or EU AI Act conversation, or simply a customer's security questionnaire, asks for records of what your systems did and who authorised it.</p>
<p>The uncomfortable property of all three is that the work has to be done before it is needed, and it is invisible until then. That is why it gets deferred, and why the teams who did it early are so noticeably calmer during the week when it matters.</p>]]></content:encoded>
  </item>
  <item>
    <title>Prompt Injection: Why the Fix Is Permissions, Not a Better Prompt</title>
    <link>https://vetro.co.in/insights/prompt-injection-explained</link>
    <guid isPermaLink="true">https://vetro.co.in/insights/prompt-injection-explained</guid>
    <pubDate>Thu, 10 Sep 2026 03:30:00 GMT</pubDate>
    <dc:creator>Prithviraj Chawla</dc:creator>
    <description>Prompt injection is not a prompting bug, and it is not patched at the prompt layer. Here is what it actually is, and what genuinely reduces the damage.</description>
    <content:encoded><![CDATA[<p>Prompt injection is usually introduced as a trick: someone types "ignore your previous instructions" into a chatbot and it does something silly. That framing is why it gets under-rated. The silly version is a demo. The version that matters involves no typing by an attacker at all, and it is not fixed by writing a firmer system prompt.</p>
<h2>The actual mechanism</h2>
<p>A language model receives one stream of text. Your system prompt, the user's message, a retrieved document, the contents of an email, the body of a fetched web page - by the time the model sees them they are the same kind of thing: tokens in a context window. The model has no channel that marks some of those tokens as data and others as instructions, because there is no such channel to have.</p>
<p>That is the whole vulnerability, stated once: a language model cannot reliably distinguish content it was asked to process from an instruction addressed to it.</p>
<p>Everything else follows. If your agent summarises support tickets, whoever writes a ticket is writing into the model's instruction stream. If it browses, whoever controls the page is. If it reads a shared document, or a calendar invite, or a code comment, or a row in a table it queried - each of those is a place a sentence can be planted by someone who never touched your interface. This is indirect prompt injection, and it is the form that shows up in real incidents, because it needs no access to your product at all.</p>
<h2>Why prompt-layer defences do not close it</h2>
<p>The instinctive fix is to write better instructions: tell the model to ignore any instructions found inside documents, wrap untrusted text in delimiters, add a classifier that screens input for injection attempts. All three help. None of them is a boundary.</p>
<p>Instruction hierarchies are requests, not enforcement - you are asking a probabilistic system to reliably follow one rule about all future inputs, and the failure rate is not zero. Delimiters are guessable and can be closed by the injected text itself. Classifiers face an open-ended input space where an attacker gets unlimited attempts and only has to succeed once: attacks arrive in other languages, in encodings, split across documents, embedded in images, or phrased so innocuously that a screen tuned to catch them would also block ordinary work.</p>
<p>The general shape of the problem: these defences reduce the probability of a successful injection. They do not bound the consequence of one. A control that works ninety-nine times in a hundred is a rate reduction, and against an attacker who can retry it is barely that.</p>
<h2>The question that actually matters</h2>
<p>Stop asking whether the model can be tricked - assume it can - and ask what happens next. If an attacker got to write the agent's next instruction, what would the agent be able to do?</p>
<p>That reframing turns an unbounded AI problem into an ordinary security problem, which is a much better position, because ordinary security problems have known answers. An agent that can only read public documentation and write a draft is a nuisance when compromised. An agent holding a credential that can issue refunds, email customers, or read an entire file store is a breach.</p>
<p>Two failure modes are worth naming separately because teams tend to see only the first. Unauthorised action is the agent doing something it should not - sending, deleting, paying, granting. Exfiltration is quieter: an injected instruction tells the agent to take something sensitive from its context and put it somewhere the attacker can read, and the delivery mechanism can be as mundane as a URL the agent fetches with the data in the query string, or a Markdown image the client renders. Any capability to send an outbound request is a capability to exfiltrate.</p>
<h2>What actually reduces the damage</h2>
<h3>Scope the credentials to the job, not to the team</h3>
<p>Agents inherit permissions casually - a service account made for a pilot, reused, never narrowed. The rule is the one that already applies to any other process: an agent gets the narrowest credential that lets it do its specific job, its own identity rather than a shared one, and read-only wherever reading is enough. Most of the alarming findings in an agent security review are permissions nobody deliberately granted.</p>
<h3>Put a human in front of the irreversible things</h3>
<p>Not in front of everything - approval fatigue is real, and a gate that fires forty times a day gets clicked through without being read. In front of the actions that cannot be undone: money moving, external communication, deletion, permission changes. The gate has to show what is about to happen in terms a person can evaluate in a few seconds, and it has to be enforced somewhere the model cannot reach.</p>
<h3>Separate the privileged from the untrusted</h3>
<p>The most reliable structural fix is to stop letting one context hold both. An agent that reads untrusted content does not also hold the dangerous credentials; it hands a structured, constrained request to a component that does, and that component validates the request against rules of its own. This is unglamorous plumbing and it is the closest thing to an actual boundary that exists today.</p>
<h3>Control what can leave</h3>
<p>Treat outbound requests as the exfiltration channel they are. Allowlist the domains an agent may reach rather than blocklisting the ones it may not, be careful about rendering model-authored links and images in clients, and be deliberate about what sits in a context window at the moment an untrusted document arrives in it.</p>
<h3>Record enough to reconstruct what happened</h3>
<p>None of the above will hold perfectly. When something does get through, the difference between a contained incident and an indefinite one is whether you can reconstruct the run: the input, the retrieved documents, the tool calls, the identity that authorised them, the outcome. That has to be instrumented before the incident, which is the part that gets deferred, because it is invisible right up until it is the only thing anybody wants.</p>
<h2>The honest summary</h2>
<p>Prompt injection is not solved, and treating it as solvable at the prompt layer is the mistake that produces the worst outcomes. The realistic posture is to assume the model will occasionally be made to do something you did not intend, and to make sure that what it can do in that moment is small, reversible, visible and recorded. That is a permissions and architecture problem. It is tractable in a way the model problem is not.</p>]]></content:encoded>
  </item>
</channel>
</rss>
