Error monitoring for the agent era
Error monitoring that proves the fix.
Crashes, releases and agent failures in one issue inbox, on ingest that speaks the Sentry envelope protocol — change the DSN and keep the SDK you already ship.
Then the part the category stops at. Your agent says it fixed the bug; Bucker reproduces the failure, runs the patch in a sandbox, and hands you an artifact you can check without trusting us. An error essence under 500 tokens. The approval inside Claude Code. A bill that cannot spike. And a human on the merge button — structurally, not as a setting.
Sentry-compatible ingest · OTLP · change the DSN, keep the SDK
- kill_switchAI on · PR creation onok
- essenceCHECKOUT-42 projected396 / 500 tokens
- repro_command_supportedpnpm vitest run · nodeok
- repro on HEADTypeError named in outputFAIL · REPRODUCED
- patch2 files · +14 −3 · in blast radiusapplied, zero fuzz
- repro post-patchsame test, same digestPASS
- suitepnpm vitest run212 passed · 0 failed
- certificateissue → repro → patch → suite → tiersigned
- approvalpr.propose is SENSITIVEwaiting for a human
- estimated tokens per error essence, count published
- ≤ 500
- MCP tools, approval-on-the-wire across every write
- 29
- containers to self-host, and a test that fails at five
- 4
- SDKs: browser, Node, Python, Flutter (Dart errors; no native crash capture)
- 4
- auto-merges, at any tier — by design; there is no merge method to call
- none
Where Bucker actually stands
Pre-launch, and specific about it
Bucker is not generally available, and this is the top of the page rather than the bottom of it. The backend is substantially ahead of the surface: the ingest pipeline, the essence contract, the remediation loop, agent governance, the proof ledger and the MCP server are implemented and tested, and the web application now reaches most of them. What is not done is named on the page where it comes up rather than in a footnote — settlement has never charged a card, the hosted-sandbox provider has never been run against the real service, and no independent witness countersigns the transparency log.
The changelog records what shipped in each phase, including the parts that did not. The documentation is rendered straight from the repository, so it describes the same product the code does.
Install
Change the DSN. Keep the SDK.
Ingest speaks the Sentry envelope protocol, so the SDK you already ship keeps working when you point it here. The first-party SDKs add what an automated fix needs — source context on the server, a pruned semantic selector on every click, a Flutter symbol upload that refuses a build whose shipped binary has no matching symbol file.
Already on another vendor?
The landing lane pulls your existing issues out of Sentry, Datadog and CloudWatch and reconciles them against a fidelity report, so the day-one inbox is not empty.
Or let the agent install it
npx create-bucker --emit-agent-docswritesllms.txt,AGENTS.mdand a checklist with explicit stop conditions, and a coding agent takes it from there with no human in the loop.What Flutter does not capture yet
The Flutter SDK captures Dart-layer errors and uploads symbols. It captures no native crashes — no Android NDK, no
PLCrashReporter, no ANR detection, no iOS MetricKit hangs — so a SIGSEGV in a plugin is invisible to it. That is most of the crash volume on a real mobile app, the seam says so in its own source, and there is no separate mobile app to evaluate. Do not choose Bucker for mobile today.Small on purpose
The browser SDK is about 11 KB gzipped, with the budget enforced in CI. The Python SDK has zero runtime dependencies, because every one it added would be one you inherit.
The error essence
What your agent reads
Every incumbent’s MCP server retrofits human-shaped JSON: unbounded payloads, whole stack traces, breadcrumb arrays designed for a scrollbar. That is your context window and your agent bill. Bucker projects every issue to a bounded document, publishes the token count on the response, versions the schema so your prompts do not break underneath you, and labels every string an attacker could have written.
{
"event_id": "b1c3…",
"exception": { "values": [ { "type": "TypeError",
"value": "Cannot read properties of undefined (reading 'price')",
"stacktrace": { "frames": [ … 22 frames, node_modules included … ] } } ] },
"breadcrumbs": { "values": [ … 40 entries, every fetch and click … ] },
"contexts": { "browser": …, "os": …, "device": …, "trace": … },
"request": { "headers": { … 31 headers … } },
"tags": { … }, "extra": { … }, "sdk": { … }, "modules": { … 140 packages … }
}- title
- TypeError: Cannot read properties of undefined trusted
- culprit
- app/checkout/total.ts · computeTotal
- exceptionValue
- …(reading 'price') untrusted
- frames[0]
- total.ts:42:11 · inApp · symbolicated
const lines = cart.items > const total = lines.reduce((s, l) => s + l.price, 0) return applyDiscount(total, cart.coupon)
- hiddenFrameCount
- 17
- blastRadius
- 5,120 events · 311 users · production · first seen 94 s after deploy 1.4.2
- suspectCommits
- a91f2c0 “checkout: coupon lines carry no price” code
- hydration
- hydrate_frames · available · 940 tokens · 17 of 22 frames omittedhydrate_repo_context · available · 1,800 tokenshydrate_locals · unavailable · local-variable capture is opt-in
- securityNotice
- null quarantinedwould replace any string the injection scanner flagged
The format is published as a standalone specification with a JSON Schema generated from the same source the server uses. The contract and the specification.
The loop
The agent is the first responder. The human is the approver.
Six stages, in this order, because each one is unreachable until the previous one holds.
Capture
Any Sentry-compatible SDK, OTLP, or your existing vendor
Envelopes, traces and logs land in one stream. A landing lane pulls issues out of Sentry, Datadog and CloudWatch and reconciles them with a fidelity report.
Compress
One bounded essence per issue
Under 500 estimated tokens, the count on every response, a reduction ladder that reports what it dropped, and handles that price the rest before you pay for it.
Explain
Ranked hypotheses that cite their evidence
Root-cause analysis phrases facts it was handed rather than deciding truth. Each hypothesis carries a confidence value and the signal it rests on — or the verdict is INCONCLUSIVE, said out loud.
Reproduce
The repro must fail first, for the right reason
A failing test on HEAD is qualified as REPRODUCED, HARNESS_ERROR or UNRELATED_FAILURE. Only the first proceeds, so a run that proves nothing also spends nothing.
Prove
A tier that is achieved, never attempted
Tier A needs the full suite green, B the affected tests, C ran no suite and is not billable. The result travels on a signed certificate that refuses to attest a checkout that never happened.
Approve
A human merges. Always.
Refused to non-human principals before policy is read, filtered out of every agent token, pinned in the trust ladder — and there is no merge method in the source-control client to call.
Sandbox-verified fixes
“Verified” has a definition here
Sentry Seer, Rollbar Resolve and Datadog’s Bits all stop at “opens a pull request.” An unverified patch is a cost imposed on a reviewer. A Bucker proposal arrives with a repro that failed before the patch and passed after it, the suite result, a tier an audit function refused to inflate — and a bundle an outsider can check without trusting us.
| Claim | Evidence | Status |
|---|---|---|
| repro_fails_on_head | exit 1 · qualification REPRODUCED · test sha256 e3b0… | failed as required |
| repro_passes_post_patch | exit 0 · same test digest · patch sha256 9f86… | holds |
| suite_green | pnpm vitest run · exit 0 · 212 tests | holds |
| blast_radius_containment | 2 paths touched · 2 permitted · 0 violations | holds |
| approver_identity | alice@acme · via Claude Code · input_required | human |
| durability_window | 30 days · regression credits the invoice back | open |
# no account, no network call, a key you obtained yourself
npx bucker-verify bundle.json --key issuer.pem --explain- provenance certificate · Ed25519
- Chains issue → repro hash → patch hash → suite → tier → approver, and refuses a tier-A claim with no suite result even when the signature is valid. The loop’s own record of one run.
- proof ledger bundle · Ed25519 over DSSE
- Every claim re-derivable from the document, an RFC 6962 inclusion proof, and a signature checked against a key you supply. The one an auditor can hold.
- A
- repro reproduced, patch fixes it, full suite green · billable
- B
- same, with the affected tests green · billable
- C
- repro and patch only, no suite ran · not billable
- NONE
- no verification claim is made
Next
Point something at it
One DSN to a first event, by hand or by a coding agent. The Free plan needs no card.
MCP · input_required
Approve where you already are
The July 2026 MCP revision made input_required the standards-track human-approval primitive. Bucker’s server implements it across its write tools and renders the card as an MCP App when the client can show one. You answer the prompt in your editor; you never open a dashboard to find the item.
Approve: run the verified-fix loop on CHECKOUT-42
Lets Bucker reproduce the failure on HEAD, author a patch inside the blast radius, re-run the repro, run the suite and sign a certificate. It does not approve or merge the resulting patch.
- acting
- Claude Code, for alice@acme, under incident #82
- policy
- pr.propose is SENSITIVE · 1 human approval required
- expires
- in 15 minutes
- evidence
- essence · blast radius · 2 ranked hypotheses, each with its n
# hosted server, one line, token from your environment
claude mcp add --transport http bucker https://mcp.bucker.io \
--header "Authorization: Bearer ${BUCKER_ACCESS_TOKEN}"- issue_essence — one issue fully explained inside the 500-token ceiling, with hydration handles priced before use.
- search — compiles a plain-language request into the query grammar and publishes the compilation, so a mis-parse is visible.
- start_remediation — a durable task handle, gated by the card on the left.
- approve*, *merge* — refused at load time. A tool with either in its name cannot be registered.
Agentic observability
An agent that lied is an issue, not a trace
A retry loop against a stale replica, a tool call that returned the wrong shape, a response truncated with no error: every span in every one of those sessions reported OK. Bucker’s detectors call the same ingest function your crashes do, so the failure gets a fingerprint, a lifecycle, an alert and a remediation run — in the inbox you already triage.
| issue | where | events | users | status | ai |
|---|---|---|---|---|---|
| CHECKOUT-42TypeError: Cannot read properties of undefined | app/checkout/total.ts | 5,120 | 311 | ONGOING | fix · tier A · awaiting approval |
| AGENT-7AgentRetryLoop · lookup_order × 3, same stale replica | agent/retry_loop | 1,204 | — | REGRESSED · prompt v14 | rca · ready · 2 hypotheses |
| AGENT-8AgentToolMisuse · output_format_violation | agent/tool_misuse | 96 | — | NEW | rca · queued |
| AGENT-9AgentSilentTruncation · 4,096 of 11,020 chars returned | agent/silent_truncation | 41 | — | NEW | — |
Agent governance
Agents are principals, not a toggle
What the category ships is an organisation-level AI switch and a setting for whether the bot may open a pull request. An enterprise asking “what could this agent do, on whose authority, and who can prove it” deserves an answer with a schema behind it.
- alice@acme · sponsor
- claude-code-01 · agent
- incident #82 · blast-radius token
- delegation
- RFC 8693 token exchange · audience-bound · the presented token is the ceiling
- budget
- $12.40 of $40.00 for this incident · tightest of (agent, human, incident) wins · reserved before each generation
- scope
- issues:read · essence:read · remediation:start ·
fix:approve·pr.merge - provisioned
- SCIM 2.0 · /scim/v2/Agents · the same sync as your employees
- revocation
- automatic on a high-severity anomaly · suspended outright on two
Read-only by default, promotion offered
A capability tier is an explicit rung. Promotion is offered on a streak, never taken; demotion is automatic when the streak breaks.
Budgets that bite mid-flight
The metered model provider reserves an estimate before each generation, so a run cannot overshoot by the length of one completion.
A kill switch before any policy is read
AI organisation-wide, and agent pull-request creation separately. Both checked first, so no policy row can re-enable them.
Billing that cannot surprise
Your worst day is on us
The defining resentment in this category is not price, it is unpredictability: one bad deploy tripling the invoice, credits that expire on the first, a seat charge for the AI. Monitoring here is flat. Only the AI lane is metered, and only on an artifact you could audit.
The illustrative values behind the chart
| hour | 00 | 01 | 02 | 03 | 04 | 05 | 06 | 07 | 08 | 09 | 10 | 11 | 12 | 13 | 14 | 15 | 16 | 17 | 18 | 19 | 20 | 21 | 22 | 23 |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| events (k) | 2.1 | 1.8 | 1.6 | 1.5 | 1.7 | 2.4 | 3.9 | 5.2 | 6.1 | 412 | 1040 | 688 | 130 | 22 | 8.3 | 6.9 | 6.4 | 6 | 5.7 | 5.1 | 4.2 | 3.4 | 2.8 | 2.3 |
| fixes | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 |
Free
$0/mo
- errors
- 50,000/mo
- verified fixes
- 1, then $30
- human seats
- unlimited
- single sign-on
- not on Free
Pro
$99/mo
- errors
- 500,000/mo
- verified fixes
- 10, then $25
- human seats
- unlimited
- single sign-on
- included
Team
$349/mo
- errors
- 2.5 million/mo
- verified fixes
- 40, then $20
- human seats
- unlimited
- single sign-on
- included
Enterprise
Quoted
- errors
- 10 million/mo
- verified fixes
- 150, then $15
- human seats
- unlimited
- single sign-on
- included
Read from the same endpoint the product reads when it enforces your quota. Credits roll over, capped at three months of your allotment. Every detail, including what is not free and why, is on the pricing page.
Next
Point something at it
One DSN to a first event, by hand or by a coding agent. The Free plan needs no card.
Against the category
Against the category, with our own gaps labelled
The position in one phrase: the neutral verifier of record. “Has AI” is table stakes, and the question is what you are being asked to believe. The ledger states the category norm beside what Bucker does instead; the five claims under it are the ones nobody else is making, and the label on each is the result of checking it against the code, not a hedge.
| Dimension | The category norm | Bucker |
|---|---|---|
| What “verified” means | A model asserts the fix is correct, and a pull request appears. | A repro that failed on HEAD for the right reason and passed after the patch, a suite result, and a tier an audit function refuses to inflate. |
| Whether you can check it | Verification is something the vendor did and reports on. | A signed bundle that bucker-verify re-derives offline, against a key you supply — no account, no network call. |
| Where you approve | A dashboard. MCP servers, where they exist, are read surfaces. | Inside Claude Code, Cursor or VS Code, through input_required on the wire. |
| What the agent reads | Human-shaped JSON: whole stack, unbounded breadcrumbs, no published size. | Under 500 estimated tokens, count published, schema pinnable, every string trust-labelled. |
| AI application failures | Traces and scores in a separate product with its own inbox. | The same issue stream, fingerprint, lifecycle, alerting and remediation loop as a crash. |
| The AI lane, priced | A per-contributor seat for the debugging agent, on top of the plan. | Unlimited human seats. Metered only on a tier A or B fix a named human approved. |
| A spike in traffic | Metered per event, so an incident is also an invoice. | Over-quota events cost nothing, degrade to stratified sampling, and are receipted. |
The left column is a category norm, not any particular vendor. The full comparison, including where established products are genuinely ahead of a pre-launch one, is on the compare page.
Five things nobody else is doing
Each one labelled with what is not true yet. A vendor that labels its own gaps is the one whose “verified” you can believe.
- claim · 01Partly shipped
Verifiable “verified”
Outcome pricing has a known fatal flaw: resolution definitions can be gamed. So a Bucker fix carries a Proof Ledger bundle — an Ed25519 signature over DSSE, an RFC 6962 Merkle inclusion proof, and every claim re-derivable from the document itself. Thebucker-verifyCLI checks one offline, makes no network calls, and can be pointed at a key you supply instead of the one embedded in the bundle.What is not true yet: A browser verifier now ships at
/verify, and a public transparency register at/transparency. Read the scope before trusting either: the browser verifier answers two questions and keeps them apart — whether every claim is supported by the evidence in the bundle, and whether we signed it — because a document that is self-consistent and a document we signed are different facts. The per-fix provenance certificate is signed with the same published key, so it is checkable too. Whether a transparency log is actually being operated is a deployment question, and the register says so plainly when none is. - claim · 02Partly shipped
Honest accuracy numbers
Competitors advertise resolution rates around 92% against a neutral baseline nearer 25%. Bucker’s scoreboard is built so a rate cannot be rendered without its denominator — the type system enforces it — and the transparency module keeps a miss-rate register rather than a highlight reel.What is not true yet: The per-organization accuracy dashboard now ships at
/verification: verified-fix rate, repro-first-fail rate, in-window regression rate and reversal debt, each beside the sample size it was computed from, and each left null rather than shown as zero when there is nothing to divide. Every figure is lifetime-scoped and labelled as such, so no rolling window is implied that the queries do not compute. It is a tenant-scoped page rather than a public one, so we are still not quoting a headline number here. - claim · 03Shipped
Approvals inside the coding agent
The July 2026 MCP revision madeinput_requiredthe standards-track human-approval primitive. Bucker’s MCP server implements it on the wire across 29 tools, and renders the approval and evidence cards as MCP Apps when a client negotiates for them. You approve a fix in Claude Code, where you already are — not by going to a dashboard. - claim · 04Shipped
Agent failures as real issues
A tool-schema violation that returned HTTP 200 is not a trace to go hunting for. Its detector calls the same ingest function your crashes do, so it becomes “Issue #412: output format violation, 1,204 occurrences, regressed in prompt v14, assigned, Ongoing” — one fingerprint, one lifecycle, one alerting path, one remediation loop. - claim · 05Partly shipped
Billing that cannot surprise
Flat platform plans with unlimited human seats, plus one metered lane that bills only for fixes a sandbox verified and a human approved. Over-quota events cost nothing — they degrade to stratified sampling and are receipted, never silently dropped. A storm from one bad deploy is one issue, so it can produce at most one billable fix.What is not true yet: No payment provider is connected, so nothing in this system has ever charged a card. The ladder, the caps and the billability rule are implemented and tested, as is the payment integration behind them. Connecting it is more than one value: the secret key alone decides only whether a provider exists, so setting it without a price id for each tier sold flips every money-shaped response to connected while a plan change still fails.