LLM observability answers “is the model behaving”, for engineers, on storage built to be cheap and short-lived. That is a good product. It is not the artefact you hand a regulator.
Both watch the same agent do the same thing. They are built to answer different people, and it shows in every design decision underneath.
The shortest-lived trace store in a typical stack keeps roughly seventy times less than the shortest obligation it would have to satisfy.
If you already emit traces, you have done most of the instrumentation work. We read the same spans.
OpenTelemetry spans, or the OpenInference instrumentors for OpenAI, Anthropic and LangChain.
Model calls become recorded actions, tool executions become tool calls, and payloads become fingerprints.
Sealed, countersigned, and checkable by someone who has never heard of your stack.
Including the one where the answer is that you may not need us yet.
Usually. They answer different questions and neither substitutes for the other. If nobody is going to ask you to prove what an agent did — no regulator, no auditor, no enterprise customer with a security review — you probably need observability and not this.
Retention is the easy half. The hard half is that a stored trace can be edited and nothing about it would show, and that the actor on it is usually a service account rather than the person who authorised the action. Longer retention gives you more of a record that was never built to be evidence.
No, and we would rather say so. Those tools are good at the job they do, and we read the same OpenTelemetry spans they do — if you already emit traces, we capture from them. We sit downstream, turning what happened into something a third party can verify.
Nothing extra. Pricing is a flat band by agent count, never per event — metering events would pay you to record less, and an incomplete record is the one failure this category cannot survive.