Skip to content

Service 02

Accuracy & Evaluation

Measuring what your system actually gets wrong, on your own data, behind a regression gate that fails the build - rather than a demo that happened to work on the day it was shown.

Most AI features ship on impressions. Somebody tried twenty prompts, the answers read well, and it went live. That is not a measurement. Twenty hand-picked prompts tell you the system can work; they tell you nothing about how often it does not, on which inputs, or whether last week's prompt edit made it quietly worse.

The work is building the unglamorous thing: a set of real inputs with known-correct outputs, drawn from your own domain and including the awkward cases people actually send. Then a harness that runs it on every change to a prompt, a model version, a retrieval index or a routing rule, and a number somebody is responsible for watching.

Grounding is where most of the error turns out to live. The failure is rarely a model inventing something from nothing; it is a model answering confidently from a document that did not contain the answer, or citing a source it did not read. That is measurable - does the claim follow from the retrieved text, yes or no - and measuring it separates a retrieval problem from a generation problem, which are fixed in completely different places.

What we will not claim: we cannot make a model correct. We can tell you how often it is not, on what kind of input, and put a gate in front of production so a regression has to get past a failing test before it reaches a user.

Questions About This Work

What does an LLM evaluation engagement produce?

An eval set of real inputs with verified correct outputs drawn from your own domain, a harness that runs it on every change, a grounding check for anything that answers from documents, an error taxonomy that says what kind of wrong and how often, and a gate in CI with a threshold agreed before the first red build.

Can you measure hallucination?

Partly, and the honest part is the useful one. For systems that answer from documents we measure grounding - whether each claim follows from the retrieved text - which splits "it hallucinates" into a retrieval problem and a generation problem that are fixed in different places. What we cannot do is certify general truthfulness, and we do not claim to.

How do you evaluate a RAG system?

Retrieval and generation separately. Retrieval: was the right document found, and did the relevant part survive truncation. Generation: given what was retrieved, is the answer supported by it, and does the system abstain when the corpus cannot answer. Scoring the two together hides which half is broken.

Do we need an eval set before going to production?

You need one before you can say anything true about the system. A hundred real cases with verified answers is about a week of labelling, and it is the difference between knowing the error rate and believing a demo. If you are already live without one, that is where the audit starts.

Start With the Audit

No engagement is scoped before a read-only review of the systems you actually run. The roadmap it produces is yours either way.

Book a Call