Skip to content

Service 05

Efficiency & Cost Control

Cost and latency measured per outcome rather than per token - model routing, caching, context discipline - and finding the calls that spend money without changing the answer.

AI spend behaves unlike most line items. It scales with usage, it arrives monthly with no unit economics attached, and very few teams can say what one completed task costs them. The invoice is a single number; the figure you actually need is cost per resolved ticket, per document processed, per outcome that mattered to somebody.

Once spend is attributed that way, the waste is usually obvious and boring. The largest available model doing work a small one does identically. Whole documents pasted into context where a paragraph would have done. Retries that silently double a call. Chains where two of the five steps change nothing downstream. Prefix caching left switched off on a prompt that is byte-identical on every single request.

The order matters more than the techniques. Measure, then route, then trim. Routing before measuring is how a team downgrades the one call that needed the larger model and finds out about it from a customer.

Efficiency is an accuracy question wearing different clothes, which is why we do not sell it on its own. Any change to routing, caching or context is a change to behaviour, so it goes through the eval set before it goes anywhere near production.

Questions About This Work

How do you reduce LLM costs without hurting quality?

Measure first, then route, then trim, and put every change through the eval set before production. The savings that survive that gate are the real ones: a smaller model where the evals say it passes, prefix caching on prompts that are byte-identical per request, context trimmed to what the answer actually used, and retries that no longer silently double a call.

What is cost per outcome, and why measure it?

The cost of one completed unit of work - a resolved ticket, a processed document, a qualified lead - rather than the monthly invoice or a per-token rate. It is the only figure that lets you compare a cheaper model that needs two attempts against a dearer one that needs one, and it is the number nobody has when the finance question arrives.

Is model routing safe?

Only after measurement. Routing a call to a smaller model is a behaviour change, and without an eval set you find out which call needed the larger model from a customer. With one, routing is a decision the numbers make for you, which is why we do not sell efficiency work on its own.

Why is efficiency bundled with accuracy?

Because every efficiency lever - routing, caching, context, retries - changes what the model sees or which model sees it, and therefore what it says. Treating cost as separate from correctness is how a team ships a cheaper system that is quietly worse.

Start With the Audit

No engagement is scoped before a read-only review of the systems you actually run. The roadmap it produces is yours either way.

Book a Call