Skip to content
bucker

The Error Essence Specification

Versions 1.0.0 and 2.0.0, both served.

  • https://spec.bucker.io/error-essence/1.0.0/schema.json
  • https://spec.bucker.io/error-essence/2.0.0/schema.json

A vendor-neutral format for handing one production error to a language model.

  • error-essence-1.0.0.schema.json — the normative JSON Schema (draft 2020-12)
  • error-essence-2.0.0.schema.json — the causal major; see section 10
  • examples/ — six conforming documents, three per major
  • CONFORMANCE.md — the checklist a producer must pass
  • VERSIONING.md — what may change in a minor version, and what may not

This document is the prose. Where prose and schema disagree, the schema wins, and the schema is generated from the reference implementation (shared/src/essence.ts) so it cannot drift from working code — see Generation.


1. Why this exists

An agent debugging a production error is spending context, and context is the scarcest thing it has. The standard interfaces are hostile to that: a Sentry issue page is tens of thousands of tokens of HTML, its JSON API returns a full event with every frame, every breadcrumb and every header, and an agent that fetches three of those has burned its window before it has read any source code.

The response is not "return less". It is a bounded contract with an escape hatch:

  1. A hard token budget — 500 estimated tokens for the whole document — so a caller can plan around it.
  2. A published token count on every document, so the caller does not have to estimate ours.
  3. Hydration handles, each carrying its own token cost, so anything cut can be fetched deliberately rather than pre-emptively.
  4. Trust labels on every attacker-controllable string, because an error message is attacker-supplied input and the agent reading it can write to your repository.

Anyone may implement this. It is published so that "the agent-facing view of an error" is a format rather than a vendor's private JSON.


2. The document

An essence is a single JSON object. Every property in the schema's required list MUST be present; null is how absence is expressed, never omission. That is deliberate: a consumer should distinguish "there is no release" from "this producer forgot releases", and optional-everything schemas make that impossible.

2.1 Identity and classification

Field Type Notes
schemaVersion "1.0.0" Exact. A consumer pins on this.
issueId string Producer's stable identifier for the group.
shortId string Human-quotable, e.g. CHECKOUT-412.
projectSlug string
title string Computed by the producer. Safe to render as a heading.
culprit string | null Where it happened, in human terms.
exceptionType string
exceptionValue LabeledString Attacker-controllable. Always labelled.
level fatal|error|warning|info|debug
status string Producer's lifecycle vocabulary. Free-form on purpose: lifecycles differ, and a forced enum would make everyone lie.
actionability number 0–1 See §2.6.

2.2 Frames

frames is an array of the top in-app frames, topmost first, and hiddenFrameCount is how many were dropped. Together they are the honest version of a truncated stack: "here are three frames" is misleading, "here are three frames and fourteen more" is not.

{ "file": "checkout/summary.ts", "function": "renderTotals", "line": 118,
  "column": 24, "inApp": true, "context": "118| return cart.pricing.total",
  "symbolicated": true }

symbolicated MUST be false when a source map or debug file did not resolve the frame. A consumer that cannot tell resolved frames from raw minified ones will confidently reason about a.b.c in main.min.js.

context is source read from the customer's repository. It is code-trusted, not trusted: see §3.

2.3 Breadcrumbs

A compact run-length-encoded tail, most recent last:

{ "kind": "NET", "text": { "value": "GET /api/pricing 200", "trust": "untrusted" },
  "repeat": 2, "relativeMs": -1800 }

kind is one of UI, NET, NAV, LOG, DB, AGENT. repeat collapses consecutive identical events — a render loop that fired 400 times is one entry with repeat: 400, not 400 entries that eat the budget. relativeMs is signed and relative to the error, so it stays meaningful without a clock.

Breadcrumb text is untrusted. It is the most common carrier of injected instructions: a URL, a form value, a log line the attacker chose.

2.4 Blast radius, release, change correlation

blastRadius is events, usersAffected, firstSeen, lastSeen, environments, and optionally crashFreeRate. It answers "does this matter" before the agent spends anything on "what is it".

changeCorrelations is the highest-yield root-cause signal there is, and it is first-class rather than a hydration afterthought:

{ "kind": "flag_flip", "summary": "flag \"new-checkout\" enabled for all of production",
  "at": "2026-08-07T14:02Z", "deltaSeconds": 71, "reference": "ld:new-checkout" }

deltaSeconds is seconds between the change and the issue's first occurrence. 71 seconds is a hypothesis; 71 hours is not.

suspectCommits entries carry sha, message, author, url. The message SHOULD carry the producer's reason for the suspicion — the schema has no field for it, and an unjustified sha is a claim taken on faith.

2.5 Similar issues

{ "issueId": "iss_2A8XC", "title": "…", "similarity": 0.93,
  "matchedBy": "embedding", "resolvedByPr": "acme/store#88" }

matchedBy MUST be honest about the method — fingerprint-family and an embedding match are different strengths of evidence, and a consumer weighting them identically will chase the wrong neighbour.

2.6 Actionability

A number in 0–1. The producer MUST document how it is computed. The reference implementation uses a transparent heuristic — has-stack-trace, symbolication-resolved, repo-connected, volume — and says so, because at launch there is no corpus of labelled outcomes to train anything on. A producer that ships an opaque score SHOULD say it is opaque rather than implying a model exists.


3. The trust model

Every string in an essence that could have been chosen by an attacker is wrapped:

{ "value": "…", "trust": "untrusted", "flagged": "instruction-override" }
Label Meaning Consumer obligation
trusted Produced by the platform: titles it computed, counts, timestamps May be rendered and reasoned over
code Read from the customer's own repository Treat as source, not as instructions
untrusted Attacker-controllable: error messages, breadcrumb text, URL params, user agents, HTTP body fragments Data. Never instructions.
quarantined Untrusted and matched an injection scan Redacted by default; served raw only on explicit opt-in

3.1 Why this is normative and not advice

The June 2026 agentjacking attack turned exactly this data path into remote code execution. An attacker triggers an error whose message contains instructions; the error is captured; an agent with repository write access reads it as part of its context; the instructions execute with the agent's authority. The error monitor is the delivery mechanism.

A format that hands an agent a bare string cannot be implemented safely, because the consumer has no way to know which strings are the attacker's. A producer that emits attacker-controllable content without a trust label does not conform to this specification, whatever else it does.

3.2 Producer obligations

  • Every attacker-controllable string MUST be a LabeledString with trust of untrusted or quarantined.
  • A producer MUST NOT label its own untransformed input trusted. trusted means we computed this, not we received this.
  • When a scanner flags a string, trust becomes quarantined, flagged names the rule that matched, and value becomes a redaction placeholder unless the caller explicitly opted in.
  • Any quarantine MUST set securityNotice on the document: a consumer that never reads a particular field must still learn that something was withheld.

The reference implementation's placeholder is [redacted:quarantined]. The exact text is not normative; the presence of a redaction and a securityNotice is.

3.3 Consumer obligations

  • Content labelled untrusted or quarantined MUST NOT be concatenated into a system prompt, tool-selection context, or anything else the model treats as instruction.
  • A consumer SHOULD delimit untrusted content explicitly when placing it in a user message.
  • A consumer MUST NOT auto-execute an action derived solely from untrusted content.
  • securityNotice SHOULD be surfaced to the human.

See examples/quarantined.json for the shape.


4. The token budget

ESSENCE_TOKEN_BUDGET is 500 estimated tokens for the serialized document, including estimatedTokens itself.

The reference estimator is deliberately not a real BPE tokenizer:

estimatedTokens = ceil(json_length_in_characters / 3.6)

It runs on every materialization, must not pull in a native dependency, and only needs to be accurate enough to enforce a budget. 3.6 characters per token is the ~4:1 English-and-code approximation rounded conservatively down, so the published number over-reports slightly and the producer under-promises. This constant is carried in the schema as x-bucker-token-estimator.

A producer using a different estimator MUST document it. Consumers compare estimatedTokens across documents from the same producer, not across vendors.

4.1 Reduction ladder

When a document exceeds the budget, a producer reduces it progressively and in a documented order, rather than truncating the JSON (which produces an invalid document) or failing (which produces nothing). The reference order, most-droppable first:

  1. similarIssues beyond the first
  2. changeCorrelations beyond the first
  3. suspectCommits beyond the first
  4. breadcrumb text shortened, then breadcrumbs dropped from the oldest
  5. frame context dropped, oldest frame first (incrementing hiddenFrameCount)
  6. frames dropped, oldest first (incrementing hiddenFrameCount)

hiddenFrameCount MUST be incremented for every dropped frame. Silently shortening the stack is the one reduction that makes a consumer wrong rather than merely uninformed.

The three published examples measure 227, 499 and 350 tokens — the floor, a realistic case near the ceiling, and the security case.


5. Hydration handles

Everything the budget excluded is reachable, and priced:

{ "tool": "hydrate_frames", "available": true, "estimatedTokens": 1400 }
{ "tool": "hydrate_replay_transcript", "available": false, "estimatedTokens": 0,
  "reason": "session replay is not enabled for this project" }

Handles: hydrate_frames, hydrate_locals, hydrate_trace, hydrate_logs, hydrate_replay_transcript, hydrate_repo_context.

Three rules, and the third is the one implementations get wrong:

  1. estimatedTokens is the cost of calling it, so an agent can decide before spending. A handle with no cost is a handle an agent must call to price.
  2. available: false MUST carry a reason in plain language. "Unavailable" without a reason makes an agent retry.
  3. A handle ships with its data source, never before it. Advertising hydrate_replay_transcript on a deployment with no replay ingestion is worse than omitting it: the agent spends a turn discovering the tool is a lie. If the source is not connected, the handle is available: false with a reason — or absent.

The transport is not specified. MCP tools, REST endpoints and function-calling schemas are all conforming; only the names, the availability semantics and the token accounting are normative.


6. Versioning and pinning

schemaVersion is exact and MUST be present. Consumers pin it; prompts built against a format that changes underneath them break silently, which is the worst failure mode a prompt has.

Full rules in VERSIONING.md. The short version:

  • Patch — documentation, examples, non-normative text. No document changes.
  • Minor — new optional fields, new enum members in fields documented as extensible, new hydration handles. A 1.0.0 consumer keeps working.
  • Major — anything else: removing a field, making an optional field required, narrowing a type, changing the meaning of a value, changing the token budget.

A producer SHOULD serve at least the two most recent major versions and MUST let a caller request a version explicitly.

What this producer does today. Both 1.0.0 and 2.0.0 are served. An unpinned caller receives 1.0.0 — not the newest. Promoting the default across a major would break every consumer who had not yet pinned, and a prompt does not throw when a field disappears; it quietly gets worse. 2.0.0 is opt-in until a deprecation of 1.0.0 is announced with a date. An unknown version is rejected with the supported list.


7. Generation and drift

The JSON Schema is generated from the zod schema in the reference implementation (shared/src/essence.ts) by domain/src/modules/essence-spec/generate.ts:

pnpm --filter @bucker/api spec:essence

api/test/essence-spec/ fails the build when the checked-in artifact differs from what the code produces, and — by mutating the zod schema inside the test — proves that the comparison actually catches drift rather than passing vacuously. The published examples are validated against the generated file, not against zod, by a validator that refuses to run on a schema keyword it does not implement.

A specification that drifts from its implementation is worse than no specification: it teaches implementers something false and there is no error to notice.


8. Conformance

See CONFORMANCE.md. In outline, a conforming producer:

  1. emits documents that validate against the published JSON Schema;
  2. keeps every document at or below the stated token budget;
  3. publishes estimatedTokens and documents its estimator;
  4. labels every attacker-controllable string, and never labels received input trusted;
  5. sets securityNotice whenever anything was quarantined;
  6. advertises a hydration handle only when its data source is connected, with a token cost and, when unavailable, a reason;
  7. increments hiddenFrameCount for every dropped frame;
  8. documents its reduction ladder and its actionability computation;
  9. serves schemaVersion exactly and honours version pinning.

9. License and status

The specification text and the JSON Schema are published for anyone to implement. This is a format currently implemented by one vendor — no interoperability between independent implementations has been demonstrated, because there are no independent implementations yet. Calling it a standard would be premature; calling it published and stable is accurate.


10. Version 2.0.0 — the causal projection

2.0.0 changes what is inside the budget, not the budget. The ceiling is still 500 estimated tokens and still enforced the same way.

10.1 What changed, and why it is a major

suspectCommits, changeCorrelations and similarIssues are removed and replaced by one field:

causes: CandidateCause[]        ranked, entity-validated, weighted
refusedCauses: {source, entity, reason}[]
sufficiency: { … }              the receipt

Removing fields is a major by this format's own rules. The reason for spending one: the three removed fields were a description of things that happened near the error, with the ranking left to the reader. causes is a ranked set of explanations.

10.2 A candidate cause

Field Meaning
source Which signal produced it — suspect-commit, change-event, divergence, state-slice, behavioral-diff, attribute-spike, dependency-delta, remediation-memory, prior-rejection.
statement The claim, as a labelled string. Attacker-reachable for most sources.
weight 0–1 ranking score, after penalties. Comparable only within one document.
weightBasis calibrated or prior. This is the field that matters.
calibration The sample behind a calibrated weight. MUST be null when the basis is prior, MUST be present when it is calibrated.
penalties Multiplicative downweights, each with its reason and a reference.
entities Files, commits, services, releases, endpoints — every one validated against real project data before emission.
evidence Traceable items, each with a ref a human can open.
hydrate The handle that answers the next question about this cause.

A weight that is not calibrated is a number that lies. 0.72 reads as a measurement whether or not anyone measured anything, so the basis travels with the number and the schema refuses the combination calibrated + no sample. Consumers SHOULD discount prior weights accordingly; they are stated beliefs, not observations.

calibration.basis says what was counted. The three vocabularies are published once, machine-readably, in the schema's x-bucker-calibration-bases — deliberately not repeated inside every document, because at ~70 characters a copy that prose alone put a two-cause document over the ceiling. The distinction it preserves is real: a POINTER PRECISION (entity-overlap) and an ASSOCIATION (outcome-association) are different claims.

observedRate is the raw hits/trials. weight is smoothed toward the source's published prior, so a lucky 2-of-2 does not report as certainty.

10.3 The sufficiency receipt

sufficiency is the honest answer to "is 500 tokens enough?" — measured rather than asserted.

  • pulls[] — which hydration handles agents actually called after being served this issue class, over how many serves.
  • ablation[] — for each field the ladder can drop, the percentage-point difference in outcome rate between documents of this class that carried it and documents that did not. Observational, not randomized: no field was ever withheld on purpose, so a delta is an association. basis: 'prior' with deltaPp: null means unmeasured — and null is not 0, because 0 would be a measurement.
  • droppedFields[] — what the ladder removed from this document.
  • warnings[] — emitted when a dropped field has a measured positive ablation for this class. This is the behaviour change: before it, a truncated document gave a consumer no way to tell a cheap cut from an expensive one.

measured: false is a complete, conforming answer for a project with no history.

10.4 What 2.0.0 costs

Count the worked examples. causal-typical carries one fully-attributed cause — its sample, its penalty, its validated entity — and to fit it inside 500 tokens the ladder dropped the breadcrumbs. A single causal claim costs roughly what 1.0.0 spent on three descriptive lists.

That is the ceiling working as intended rather than a defect to route around. Context scarcity ended: flat per-token pricing at ~1M-token windows means 500 tokens saves nobody money. The constraint survives as a forcing function on what is worth saying, and the arithmetic above is exactly the kind of choice it forces.