Skip to content
bucker

Agentic observability

Monitoring your AI application, in the same product as your crashes.

The thesis

Agents fail with HTTP 200 and plausible text. Every span is ok, no exception is raised, the latency looks fine — and the user got a confidently wrong answer, or a tool was called eleven times with the same broken arguments, or the retriever returned nothing and the model happily made something up.

Status codes are useless as the only signal, so the signal is derived from the shape of the run.

Today error tracking and LLM observability are separate products with separate on-call surfaces. Here they are one: an agent failure and a null-pointer exception land in the same issue stream, with one lifecycle, one alerting path and one remediation loop. That is enforced structurally — the detectors emit occurrences through ingestNormalizedEvent, the same function the crash pipeline calls. There is no parallel "agent incident" table with its own status vocabulary.

Agent issues

Each detector turns a failure shape into an occurrence. Because these occurrences have no stack trace, they carry an explicit [agent_issue, kind, subject] fingerprint — stable across runs, and legible in the grouping rationale rather than a hash nobody can explain.

False positives are the failure mode. A detector that fires on healthy traffic makes the whole product unusable, so every threshold is explicit, every detector has negative tests, and none of them fire on "the model said something I dislike".

Ingest

GenAI spans arrive over OTLP. The OTel GenAI semantic conventions are pre-stable and have been renamed repeatedly (gen_ai.system → gen_ai.provider.name; prompt_tokens/completion_tokens → input_tokens/output_tokens), and a large share of already-instrumented apps emit a different vocabulary entirely — OpenInference's llm.*/tool.*, OpenLLMetry's traceloop.*.

Hard-coding any one generation is a named risk in the design doc, so every field is read through an ordered candidate list (newest naming first, because when both are present the newer one is what the instrumentation meant) and normalized into one internal schema. Adding a fourth vocabulary is a change to one table.

Tiered evaluation

The tiering is an economic argument, not an architectural preference. Running an LLM judge over 100% of an agent's traffic costs more than the traffic did — so teams run judges on a 1% sample, miss the failures, and conclude evals do not work.

Tier Runs on What it is
1 100% of traffic Deterministic. Free and instant: schema-invalid tool arguments, retry loops, empty retrieval, output-format violations, truncation. The same detectors that produce agent issues — one implementation, so an eval and a production alert cannot disagree.
2 Only what tier 1 flagged Cheap deterministic scorer: token-overlap against expected output. No model, no network.
3 Only escalations LLM judge. Version-pinned, and the version is recorded on every result.

A score whose judge version is unknown is not comparable to any other score.

Judge drift

Uncalibrated LLM-as-judge is a documented failure mode with two halves that are constantly confused:

  • Meta-evaluation drift — the judge is nominally the same, but the model behind it moved (a provider updated a checkpoint, a temperature default changed). Yesterday's 0.82 and today's 0.71 were measured with different rulers.
  • Evaluator replacement ambiguity — the team upgraded the judge deliberately, and now nobody can say whether a score drop is the new judge being stricter or the system actually getting worse.

The module distinguishes them and attributes drift rather than reporting one number and letting you guess.

Cost

The customer-facing cost engine prices your AI application from your GenAI spans, rolled up by org, project, agent, feature and user, with cost anomalies raised as ordinary issues.

It is deliberately not the same code as our internal spend meter (which prices our own remediation agents against agent budgets). Different inputs, different consumers, different lifetimes — collapsing them would tie our internal budget arithmetic to a pricing table that changes whenever a vendor announces something.

Only inference spans are metered. Tool calls, retrieval, embeddings and agent orchestration spans are not billed as inference.

Prompts and trace replay

Prompt versions are tracked, and traces can be replayed against a different prompt or model version — which is how "did the prompt change break this, or did the model?" becomes answerable rather than a debate.

Content policy

Prompts and completions are the most sensitive payload the platform accepts. Capture is governed by a per-project policy rather than captured by default.

Where it lives

domain/src/modules/agent-observability/ — agent-issues, genai-ingest, evals, judge-drift, cost, prompts, trace-replay, content-policy.

Not built

  • No agent-simulation or synthetic-traffic testing.
  • No cross-tenant benchmark ("your agent vs the median agent").
  • Tier-3 judge results are recorded with their judge version, but there is no automatic re-scoring of history when a judge version changes.