Skip to content
bucker

Error monitoring for the agent era

Error monitoring that proves the fix.

Crashes, releases and agent failures in one issue inbox, on ingest that speaks the Sentry envelope protocol — change the DSN and keep the SDK you already ship.

Then the part the category stops at. Your agent says it fixed the bug; Bucker reproduces the failure, runs the patch in a sandbox, and hands you an artifact you can check without trusting us. An error essence under 500 tokens. The approval inside Claude Code. A bill that cannot spike. And a human on the merge button — structurally, not as a setting.

Sentry-compatible ingest · OTLP · change the DSN, keep the SDK

run · CHECKOUT-42 · sandbox localverifying…Tier A · awaiting a human
  1. kill_switchAI on · PR creation onok
  2. essenceCHECKOUT-42 projected396 / 500 tokens
  3. repro_command_supportedpnpm vitest run · nodeok
  4. repro on HEADTypeError named in outputFAIL · REPRODUCED
  5. patch2 files · +14 −3 · in blast radiusapplied, zero fuzz
  6. repro post-patchsame test, same digestPASS
  7. suitepnpm vitest run212 passed · 0 failed
  8. certificateissue → repro → patch → suite → tiersigned
  9. approvalpr.propose is SENSITIVEwaiting for a human
An illustrative run. The stage names, the qualification and the tier rule are the loop’s own; the counts are examples.
estimated tokens per error essence, count published
≤ 500
MCP tools, approval-on-the-wire across every write
29
containers to self-host, and a test that fails at five
4
SDKs: browser, Node, Python, Flutter (Dart errors; no native crash capture)
4
auto-merges, at any tier — by design; there is no merge method to call
none

Where Bucker actually stands

Pre-launch, and specific about it

Bucker is not generally available, and this is the top of the page rather than the bottom of it. The backend is substantially ahead of the surface: the ingest pipeline, the essence contract, the remediation loop, agent governance, the proof ledger and the MCP server are implemented and tested, and the web application now reaches most of them. What is not done is named on the page where it comes up rather than in a footnote — settlement has never charged a card, the hosted-sandbox provider has never been run against the real service, and no independent witness countersigns the transparency log.

The changelog records what shipped in each phase, including the parts that did not. The documentation is rendered straight from the repository, so it describes the same product the code does.

Install

Change the DSN. Keep the SDK.

Ingest speaks the Sentry envelope protocol, so the SDK you already ship keeps working when you point it here. The first-party SDKs add what an automated fix needs — source context on the server, a pruned semantic selector on every click, a Flutter symbol upload that refuses a build whose shipped binary has no matching symbol file.

Choose an SDK
import { init } from '@bucker/browser'

init({
  dsn: 'https://<publicKey>@us.ingest.bucker.io/<projectId>',
  environment: 'production',
  release: 'web@1.4.2',
})
import { init, requestHandler, errorHandler } from '@bucker/node'

const client = init({
  dsn: 'https://<publicKey>@us.ingest.bucker.io/<projectId>',
  release: 'api@3.1.0',
})

app.use(requestHandler(client))
app.use(errorHandler(client))
import bucker_sdk

bucker_sdk.init(
    dsn="https://PUBLIC_KEY@us.ingest.bucker.io/PROJECT_ID",
    environment="production",
    release="api@1.4.0",
)
await Bucker.init(
  dsn: 'https://<publicKey>@us.ingest.bucker.io/<projectId>',
  environment: 'production',
  release: 'com.acme.app@1.4.2+118',
);
npx create-bucker --dsn "https://<publicKey>@us.ingest.bucker.io/<projectId>"

# detects the package manager and framework, installs the SDK,
# writes the init files and the DSN into an env file, and adds
# a /bucker-verify route that throws on purpose.
  • Already on another vendor?

    The landing lane pulls your existing issues out of Sentry, Datadog and CloudWatch and reconciles them against a fidelity report, so the day-one inbox is not empty.

  • Or let the agent install it

    npx create-bucker --emit-agent-docs writes llms.txt, AGENTS.md and a checklist with explicit stop conditions, and a coding agent takes it from there with no human in the loop.

  • What Flutter does not capture yet

    The Flutter SDK captures Dart-layer errors and uploads symbols. It captures no native crashes — no Android NDK, no PLCrashReporter, no ANR detection, no iOS MetricKit hangs — so a SIGSEGV in a plugin is invisible to it. That is most of the crash volume on a real mobile app, the seam says so in its own source, and there is no separate mobile app to evaluate. Do not choose Bucker for mobile today.

  • Small on purpose

    The browser SDK is about 11 KB gzipped, with the budget enforced in CI. The Python SDK has zero runtime dependencies, because every one it added would be one you inherit.

The error essence

What your agent reads

Every incumbent’s MCP server retrofits human-shaped JSON: unbounded payloads, whole stack traces, breadcrumb arrays designed for a scrollbar. That is your context window and your agent bill. Bucker projects every issue to a bounded document, publishes the token count on the response, versions the schema so your prompts do not break underneath you, and labels every string an attacker could have written.

the category · one event, human-shapedsize unpublished
{
  "event_id": "b1c3…",
  "exception": { "values": [ { "type": "TypeError",
    "value": "Cannot read properties of undefined (reading 'price')",
    "stacktrace": { "frames": [ … 22 frames, node_modules included … ] } } ] },
  "breadcrumbs": { "values": [ … 40 entries, every fetch and click … ] },
  "contexts": { "browser": …, "os": …, "device": …, "trace": … },
  "request": { "headers": { … 31 headers … } },
  "tags": { … }, "extra": { … }, "sdk": { … }, "modules": { … 140 packages … }
}
Thousands of tokens, most of them frames an agent will never read, and no way to know the size before spending it.
issue_essence · CHECKOUT-42 · schemaVersion 1.0.0estimatedTokens 396
budget396 of 500 · target 400
title
TypeError: Cannot read properties of undefined trusted
culprit
app/checkout/total.ts · computeTotal
exceptionValue
…(reading 'price') untrusted
frames[0]
total.ts:42:11 · inApp · symbolicated
  const lines = cart.items
> const total = lines.reduce((s, l) => s + l.price, 0)
  return applyDiscount(total, cart.coupon)
hiddenFrameCount
17
blastRadius
5,120 events · 311 users · production · first seen 94 s after deploy 1.4.2
suspectCommits
a91f2c0 “checkout: coupon lines carry no price” code
hydration
hydrate_frames · available · 940 tokens · 17 of 22 frames omitted
hydrate_repo_context · available · 1,800 tokens
hydrate_locals · unavailable · local-variable capture is opt-in
securityNotice
null quarantinedwould replace any string the injection scanner flagged
The count rounds against us, the ladder reports what it dropped, and every string an attacker could write carries its trust label.

The format is published as a standalone specification with a JSON Schema generated from the same source the server uses. The contract and the specification.

The loop

The agent is the first responder. The human is the approver.

Six stages, in this order, because each one is unreachable until the previous one holds.

  1. Capture

    Any Sentry-compatible SDK, OTLP, or your existing vendor

    Envelopes, traces and logs land in one stream. A landing lane pulls issues out of Sentry, Datadog and CloudWatch and reconciles them with a fidelity report.

  2. Compress

    One bounded essence per issue

    Under 500 estimated tokens, the count on every response, a reduction ladder that reports what it dropped, and handles that price the rest before you pay for it.

  3. Explain

    Ranked hypotheses that cite their evidence

    Root-cause analysis phrases facts it was handed rather than deciding truth. Each hypothesis carries a confidence value and the signal it rests on — or the verdict is INCONCLUSIVE, said out loud.

  4. Reproduce

    The repro must fail first, for the right reason

    A failing test on HEAD is qualified as REPRODUCED, HARNESS_ERROR or UNRELATED_FAILURE. Only the first proceeds, so a run that proves nothing also spends nothing.

  5. Prove

    A tier that is achieved, never attempted

    Tier A needs the full suite green, B the affected tests, C ran no suite and is not billable. The result travels on a signed certificate that refuses to attest a checkout that never happened.

  6. Approve

    A human merges. Always.

    Refused to non-human principals before policy is read, filtered out of every agent token, pinned in the trust ladder — and there is no merge method in the source-control client to call.

Sandbox-verified fixes

“Verified” has a definition here

Sentry Seer, Rollbar Resolve and Datadog’s Bits all stop at “opens a pull request.” An unverified patch is a cost imposed on a reviewer. A Bucker proposal arrives with a repro that failed before the patch and passed after it, the suite result, a tier an audit function refused to inflate — and a bundle an outsider can check without trusting us.

proof ledger bundle · proposal prp_4f2a · CHECKOUT-42tier A
Claims in an example Proof Ledger bundle
ClaimEvidenceStatus
repro_fails_on_headexit 1 · qualification REPRODUCED · test sha256 e3b0…failed as required
repro_passes_post_patchexit 0 · same test digest · patch sha256 9f86…holds
suite_greenpnpm vitest run · exit 0 · 212 testsholds
blast_radius_containment2 paths touched · 2 permitted · 0 violationsholds
approver_identityalice@acme · via Claude Code · input_requiredhuman
durability_window30 days · regression credits the invoice backopen
# no account, no network call, a key you obtained yourself
npx bucker-verify bundle.json --key issuer.pem --explain
Example values. The claim kinds, the qualification and the “failed as required” rule are from the published Proof Ledger specification.
two signatures, two meanings
provenance certificate · Ed25519
Chains issue → repro hash → patch hash → suite → tier → approver, and refuses a tier-A claim with no suite result even when the signature is valid. The loop’s own record of one run.
proof ledger bundle · Ed25519 over DSSE
Every claim re-derivable from the document, an RFC 6962 inclusion proof, and a signature checked against a key you supply. The one an auditor can hold.
tiers are achieved, never attempted
A
repro reproduced, patch fixes it, full suite green · billable
B
same, with the affected tests green · billable
C
repro and patch only, no suite ran · not billable
NONE
no verification claim is made

Next

Point something at it

One DSN to a first event, by hand or by a coding agent. The Free plan needs no card.

MCP · input_required

Approve where you already are

The July 2026 MCP revision made input_required the standards-track human-approval primitive. Bucker’s server implements it across its write tools and renders the card as an MCP App when the client can show one. You answer the prompt in your editor; you never open a dashboard to find the item.

claude code · tools/call start_remediationinput_required

Approve: run the verified-fix loop on CHECKOUT-42

Lets Bucker reproduce the failure on HEAD, author a patch inside the blast radius, re-run the repro, run the suite and sign a certificate. It does not approve or merge the resulting patch.

acting
Claude Code, for alice@acme, under incident #82
policy
pr.propose is SENSITIVE · 1 human approval required
expires
in 15 minutes
evidence
essence · blast radius · 2 ranked hypotheses, each with its n
Rendered as an MCP App when the client negotiates for it; otherwise the same decision as structured text. Approving the run is not approving the fix.
# hosted server, one line, token from your environment
claude mcp add --transport http bucker https://mcp.bucker.io \
  --header "Authorization: Bearer ${BUCKER_ACCESS_TOKEN}"
what the 29 tools will and will not do
  • issue_essence — one issue fully explained inside the 500-token ceiling, with hydration handles priced before use.
  • search — compiles a plain-language request into the query grammar and publishes the compilation, so a mis-parse is visible.
  • start_remediation — a durable task handle, gated by the card on the left.
  • approve*, *merge* — refused at load time. A tool with either in its name cannot be registered.

Agentic observability

An agent that lied is an issue, not a trace

A retry loop against a stale replica, a tool call that returned the wrong shape, a response truncated with no error: every span in every one of those sessions reported OK. Bucker’s detectors call the same ingest function your crashes do, so the failure gets a fingerprint, a lifecycle, an alert and a remediation run — in the inbox you already triage.

issues · checkout · production · is:unresolvederrorSpanCount = 0 on every agent row
An issue stream in which agent failures are ordinary issues
issuewhereeventsusersstatusai
CHECKOUT-42TypeError: Cannot read properties of undefinedapp/checkout/total.ts5,120311ONGOINGfix · tier A · awaiting approval
AGENT-7AgentRetryLoop · lookup_order × 3, same stale replicaagent/retry_loop1,204—REGRESSED · prompt v14rca · ready · 2 hypotheses
AGENT-8AgentToolMisuse · output_format_violationagent/tool_misuse96—NEWrca · queued
AGENT-9AgentSilentTruncation · 4,096 of 11,020 chars returnedagent/silent_truncation41—NEW—
Six failing sessions became three rows. Twenty would still be three: the fingerprint is the kind and the subject, never the run.

Agent governance

Agents are principals, not a toggle

What the category ships is an organisation-level AI switch and a setting for whether the bot may open a pull request. An enterprise asking “what could this agent do, on whose authority, and who can prove it” deserves an answer with a schema behind it.

principal · agent · claude-code-01active · baseline learned
  1. alice@acme · sponsor
  2. claude-code-01 · agent
  3. incident #82 · blast-radius token
delegation
RFC 8693 token exchange · audience-bound · the presented token is the ceiling
budget
$12.40 of $40.00 for this incident · tightest of (agent, human, incident) wins · reserved before each generation
scope
issues:read · essence:read · remediation:start · fix:approve · pr.merge
provisioned
SCIM 2.0 · /scim/v2/Agents · the same sync as your employees
revocation
automatic on a high-severity anomaly · suspended outright on two
The audit row reads “Claude Code, acting for alice@acme under incident #82”, not a service-account name. Example values.
  • Read-only by default, promotion offered

    A capability tier is an explicit rung. Promotion is offered on a streak, never taken; demotion is automatic when the streak breaks.

  • Budgets that bite mid-flight

    The metered model provider reserves an estimate before each generation, so a run cannot overshoot by the length of one completion.

  • A kill switch before any policy is read

    AI organisation-wide, and agent pull-request creation separately. Both checked first, so no policy row can re-enable them.

Billing that cannot surprise

Your worst day is on us

The defining resentment in this category is not price, it is unpredictability: one bad deploy tripling the invoice, credits that expire on the first, a seat charge for the AI. Monitoring here is flat. Only the AI lane is metered, and only on an artifact you could audit.

one bad deploy · 09:00 · 24 hoursevent overage: $0
events per hourpeak 1,040,000 · one fingerprint
billable verified fixesat most 1 · tier A, human-approved
The illustrative values behind the chart
Events per hour and billable fixes, by hour
hour000102030405060708091011121314151617181920212223
events (k)2.11.81.61.51.72.43.95.26.14121040688130228.36.96.465.75.14.23.42.82.3
fixes000000000001111111111111
Past quota, ingest degrades to deterministic stratified sampling and every degraded event is receipted. The storm is one issue, so the AI lane can bill for at most one fix.
  • Free

    $0/mo

    errors
    50,000/mo
    verified fixes
    1, then $30
    human seats
    unlimited
    single sign-on
    not on Free
  • Pro

    $99/mo

    errors
    500,000/mo
    verified fixes
    10, then $25
    human seats
    unlimited
    single sign-on
    included
  • Team

    $349/mo

    errors
    2.5 million/mo
    verified fixes
    40, then $20
    human seats
    unlimited
    single sign-on
    included
  • Enterprise

    Quoted

    errors
    10 million/mo
    verified fixes
    150, then $15
    human seats
    unlimited
    single sign-on
    included

Read from the same endpoint the product reads when it enforces your quota. Credits roll over, capped at three months of your allotment. Every detail, including what is not free and why, is on the pricing page.

Next

Point something at it

One DSN to a first event, by hand or by a coding agent. The Free plan needs no card.

Against the category

Against the category, with our own gaps labelled

The position in one phrase: the neutral verifier of record. “Has AI” is table stakes, and the question is what you are being asked to believe. The ledger states the category norm beside what Bucker does instead; the five claims under it are the ones nobody else is making, and the label on each is the result of checking it against the code, not a hedge.

Bucker against the category norm
DimensionThe category normBucker
What “verified” meansA model asserts the fix is correct, and a pull request appears.A repro that failed on HEAD for the right reason and passed after the patch, a suite result, and a tier an audit function refuses to inflate.
Whether you can check itVerification is something the vendor did and reports on.A signed bundle that bucker-verify re-derives offline, against a key you supply — no account, no network call.
Where you approveA dashboard. MCP servers, where they exist, are read surfaces.Inside Claude Code, Cursor or VS Code, through input_required on the wire.
What the agent readsHuman-shaped JSON: whole stack, unbounded breadcrumbs, no published size.Under 500 estimated tokens, count published, schema pinnable, every string trust-labelled.
AI application failuresTraces and scores in a separate product with its own inbox.The same issue stream, fingerprint, lifecycle, alerting and remediation loop as a crash.
The AI lane, pricedA per-contributor seat for the debugging agent, on top of the plan.Unlimited human seats. Metered only on a tier A or B fix a named human approved.
A spike in trafficMetered per event, so an incident is also an invoice.Over-quota events cost nothing, degrade to stratified sampling, and are receipted.

The left column is a category norm, not any particular vendor. The full comparison, including where established products are genuinely ahead of a pre-launch one, is on the compare page.

Five things nobody else is doing

Each one labelled with what is not true yet. A vendor that labels its own gaps is the one whose “verified” you can believe.

  • claim · 01Partly shipped

    Verifiable “verified”

    Outcome pricing has a known fatal flaw: resolution definitions can be gamed. So a Bucker fix carries a Proof Ledger bundle — an Ed25519 signature over DSSE, an RFC 6962 Merkle inclusion proof, and every claim re-derivable from the document itself. The bucker-verify CLI checks one offline, makes no network calls, and can be pointed at a key you supply instead of the one embedded in the bundle.

    What is not true yet: A browser verifier now ships at /verify, and a public transparency register at /transparency. Read the scope before trusting either: the browser verifier answers two questions and keeps them apart — whether every claim is supported by the evidence in the bundle, and whether we signed it — because a document that is self-consistent and a document we signed are different facts. The per-fix provenance certificate is signed with the same published key, so it is checkable too. Whether a transparency log is actually being operated is a deployment question, and the register says so plainly when none is.

  • claim · 02Partly shipped

    Honest accuracy numbers

    Competitors advertise resolution rates around 92% against a neutral baseline nearer 25%. Bucker’s scoreboard is built so a rate cannot be rendered without its denominator — the type system enforces it — and the transparency module keeps a miss-rate register rather than a highlight reel.

    What is not true yet: The per-organization accuracy dashboard now ships at /verification: verified-fix rate, repro-first-fail rate, in-window regression rate and reversal debt, each beside the sample size it was computed from, and each left null rather than shown as zero when there is nothing to divide. Every figure is lifetime-scoped and labelled as such, so no rolling window is implied that the queries do not compute. It is a tenant-scoped page rather than a public one, so we are still not quoting a headline number here.

  • claim · 03Shipped

    Approvals inside the coding agent

    The July 2026 MCP revision made input_required the standards-track human-approval primitive. Bucker’s MCP server implements it on the wire across 29 tools, and renders the approval and evidence cards as MCP Apps when a client negotiates for them. You approve a fix in Claude Code, where you already are — not by going to a dashboard.
  • claim · 04Shipped

    Agent failures as real issues

    A tool-schema violation that returned HTTP 200 is not a trace to go hunting for. Its detector calls the same ingest function your crashes do, so it becomes “Issue #412: output format violation, 1,204 occurrences, regressed in prompt v14, assigned, Ongoing” — one fingerprint, one lifecycle, one alerting path, one remediation loop.
  • claim · 05Partly shipped

    Billing that cannot surprise

    Flat platform plans with unlimited human seats, plus one metered lane that bills only for fixes a sandbox verified and a human approved. Over-quota events cost nothing — they degrade to stratified sampling and are receipted, never silently dropped. A storm from one bad deploy is one issue, so it can produce at most one billable fix.

    What is not true yet: No payment provider is connected, so nothing in this system has ever charged a card. The ladder, the caps and the billability rule are implemented and tested, as is the payment integration behind them. Connecting it is more than one value: the secret key alone decides only whether a provider exists, so setting it without a price id for each tier sold flips every money-shaped response to connected while a plan change still fails.