Skip to content

Insights / 8 min Read

Measuring an AI Feature: What to Check Before You Ship It

Twenty prompts that looked right is not a measurement. What an eval set needs to contain, what to measure, and how to stop regressions reaching users.

Published 2026-09-10

Almost every AI feature ships the same way. Someone builds it, tries a couple of dozen prompts, the answers read well, a demo goes smoothly, and it goes live. Then it is in production and nobody can say how often it is wrong, on what kind of input, or whether this week's prompt edit improved anything.

The gap is not rigour for its own sake. It is that the twenty prompts were written by the person who built the thing, which means they are the inputs it was implicitly designed to handle. They demonstrate that the system can work. They say nothing about how often it does.

The eval set is the whole thing

An eval set is a collection of real inputs with known-correct outputs. Everything else in evaluation is technique; this is the part that decides whether any of it means anything, and it is the part teams try hardest to skip because it is manual and dull.

Draw from real traffic, not imagination

The single highest-value move is to build the set from inputs users actually sent, once you have any. Real inputs are misspelt, under-specified, pasted from elsewhere, several questions at once, and occasionally in a language nobody planned for. Synthetic inputs written by the team are uniformly better-formed than reality, which is precisely why a system tuned on them is disappointing in production.

Weight it towards the awkward cases

A set that mirrors production's distribution will be mostly easy cases, which wastes most of the measurement on things that were never going to fail. Deliberately over-sample the edges: questions with no answer in the corpus, where the honest response is to say so; ambiguous questions with two defensible readings; inputs near a policy boundary; long inputs that stress the context window; adversarial ones. The middle of the distribution is not where you get hurt.

A hundred good cases beats a thousand careless ones

Every case needs an output somebody knowledgeable has actually confirmed is right. That is slow, and it is also the source of all the value - an eval set with unreliable labels produces confident numbers that mean nothing, which is worse than having no numbers at all. Start at a hundred, curated carefully, and grow it from real failures as they surface.

Measure grounding separately from everything else

For any system that answers from documents, the most useful single measurement is whether each claim in the answer actually follows from the retrieved text. Not whether it is true in general, and not whether it reads well - whether it is supported by what was in front of the model.

Keeping this separate matters because it splits one vague complaint - "it hallucinates" - into two different bugs with different fixes. If the right document was retrieved and the answer still is not supported by it, that is a generation problem: prompt, model, context assembly. If the right document was never retrieved, no amount of prompt work fixes it, and every hour spent there is wasted. Measure retrieval and generation separately or you will spend your effort in the wrong half.

The other measurement worth having early: on questions the corpus genuinely cannot answer, does the system say it does not know? A system that never abstains has not been shown to be accurate, only to be confident.

On using a model as the judge

Scoring answers by hand does not scale, so the standard approach is to have a model grade them. It works well enough to be worth doing, with three caveats worth knowing before you trust the number.

  • Judges have biases - towards longer answers, towards their own family's phrasing, towards the first option shown in a pairwise comparison. Randomise order and be careful comparing across model families.
  • A judge scoring a vague criterion like "helpfulness" produces a number with no defensible meaning. Narrow, checkable questions - "is every claim supported by the source text, yes or no" - are far more reliable.
  • A judge needs its own validation. Score a hundred cases by hand, compare against the judge, and measure the agreement. If you have not done this, you do not know what your eval numbers mean.

And keep a human in the loop for the cases automation genuinely cannot score. There is always a residue, and pretending otherwise is how a metric drifts away from the thing it was supposed to represent.

Wire it to the build, or it will rot

An eval suite that has to be run manually gets run before big launches and forgotten in between, which is exactly backwards - the regressions that hurt come from small changes nobody thought needed checking.

So it runs automatically on any change to a prompt, a model version, a retrieval configuration, a tool schema or a routing rule, and it fails the build when a metric drops past a threshold agreed in advance. Agreed in advance is the operative part. A threshold negotiated after a red build is not a gate, it is a conversation, and it always resolves in favour of shipping.

One deliberate exception is worth building in: expected changes. When a prompt is meant to alter behaviour, the suite should make it easy to review the differences and accept them explicitly, rather than tempting anyone to switch the gate off.

Then keep measuring in production

An eval set is a fixed sample of a world that moves. Users ask new things, the corpus changes, providers update models underneath you. Offline evaluation catches regressions you caused; it does not catch drift you did not.

The production signals that pay for themselves are cheap: refusal and error rates, latency and cost per outcome, retrieval scores trending down, and thin implicit feedback like whether a user rephrased the same question immediately. Feed real failures back into the eval set as they surface. That loop - production failure becomes a permanent test case - is what stops the same bug shipping twice, and it is worth more than any single metric on the dashboard.

The realistic minimum

If you take one thing: a hundred real cases with verified answers, a grounding check, a threshold agreed before the build turns red, and a habit of adding every production failure back to the set. That is a weekend of work and about a week of labelling, and it is the difference between knowing what your system does and believing what your demo did.

General information, not legal or professional advice about your organisation. Regulatory obligations depend on specific facts; take advice from qualified counsel before acting. See the Terms of Use.

Questions This Raises

How many test cases does an AI eval set need?

Start around a hundred, curated carefully, and grow it from real production failures. Label quality matters far more than volume - a hundred cases with outputs a knowledgeable person has verified is more useful than a thousand carelessly labelled ones, because unreliable labels produce confident numbers that mean nothing.

What is grounding evaluation, and why measure it separately?

Grounding evaluation checks whether each claim in an answer actually follows from the retrieved source text, rather than whether it reads well or is true in general. Measuring it separately splits "it hallucinates" into two distinct bugs: if the right document was retrieved and the answer still is not supported, that is a generation problem; if the document was never retrieved, no amount of prompt work will fix it.

Can we use an LLM to grade our AI outputs?

Yes, with caveats. Judges are biased towards longer answers, towards their own family's phrasing, and towards the first option in a comparison, so randomise order and be careful across model families. Narrow, checkable criteria beat vague ones like "helpfulness". And validate the judge itself against a hundred hand-scored cases - without that, you do not know what your numbers mean.

When should AI evals run?

Automatically, on every change to a prompt, model version, retrieval configuration, tool schema or routing rule - failing the build when a metric drops past a threshold agreed in advance. A suite that has to be triggered by hand gets run before big launches and skipped in between, which misses exactly the small changes that cause most regressions.

Want This Applied To Your Systems?

The initial audit applies this to the systems you run and produces a written roadmap ordered by risk.

Book a Call