Service 04
Monitoring & Incident Response
Tracing every model call and tool call so you can say what a system did, on what input, and on whose authority - built before the incident, because it cannot be added afterwards.
Most AI deployments cannot answer the first question anyone asks. A customer says your assistant told them something wrong on Tuesday. What was the input, which prompt version was live, which documents were retrieved, what did it go on to call, and who authorised that? In a lot of systems the honest answer is that nobody knows, because the only record is an application log that captured the response and discarded everything that produced it.
Instrumentation is the fix, and it has to exist before the thing it explains. That is the entire difficulty: it is work that looks unnecessary right up to the hour it becomes the only thing that matters.
We wire tracing through the model calls, the retrieval and the tool calls, so every action carries its inputs, its context, its versions and the identity that authorised it, and a run can be reconstructed end to end. On top of that sits alerting for the things worth interrupting someone about - refusal and error spikes, latency, cost, a tool being called at a rate nobody expected, output drifting away from what the eval set says normal looks like.
It is built on open standards and into the stack you already run. OpenTelemetry traces go to your existing collector. We are not trying to sell you another console you have to remember to open.
Questions About This Work
What is LLM observability?
Being able to say, for any run your system did, what the input was, which prompt and model versions were live, what was retrieved, which tools were called with what arguments, who authorised them, and what came out - from one connected trace rather than four logs correlated by timestamp. It is instrumentation, and it has to exist before the incident it explains.
Do we need to buy an LLM observability platform?
Usually not. Traces and spans in OpenTelemetry already model an agent run natively, and they land in the collector and backend you already run. We instrument into that. A dedicated console is sometimes worth adding later; it is rarely the first thing missing.
What should we alert on?
The things worth interrupting a person for: refusal and error spikes, latency and cost per outcome moving, a tool being called at a rate nobody expected, retrieval scores trending down, and output drifting from what the eval set says normal looks like. Alerts on raw token counts tend to be noise.
How do you handle sensitive data in traces?
Redact structured identifiers at capture, store document ids and hashes rather than full text, keep payloads on a short retention window while metadata and versions stay much longer, and put access controls on the trace store equal to those on the source data. Logging nothing is not a privacy strategy; it is the absence of one.
Start With the Audit
No engagement is scoped before a read-only review of the systems you actually run. The roadmap it produces is yours either way.