Skip to content
bucker

The remediation loop and verification tiers

The atomic product unit is not "an error was captured". It is a fix that was verified. Anyone can generate a plausible-looking patch in seconds; the scarce thing is evidence that the patch was checked.

The invariant

No fix PR without a failing repro test that passes after the patch, with the verification tier stated on the PR.

That is the whole product in one sentence. Everything below is machinery to make it true and to make the claim auditable.

The pre-patch failure has to fail for the right reason

"The repro exited non-zero" and "the bug reproduced" are different claims, and only the second one supports a tier. Until August 2026 the loop computed ok = !passed(result), so any non-zero exit satisfied the invariant — a missing test runner, an unresolvable import, a syntax error in the generated test, or the default node --test {file} template pointed at a Python project. Every one of those exits non-zero and none of them says anything about the defect.

The observed failure is now qualified against the issue's signature (exception type, message shape, and the file/function anchors from its fingerprint) and classified as exactly one of:

Outcome Means Proceeds?
REPRODUCED The failure is attributable to this issue — the output names the exception, its message shape or a frame from the stack, or the repro is anchored to that code and provably ran and failed Yes
HARNESS_ERROR The test never ran: runner missing, import unresolvable, syntax error, unknown flag, no tests matched, timeout, silence No
UNRELATED_FAILURE Something failed, but nothing ties it to this issue No
DID_NOT_FAIL The repro passed on HEAD, so it reproduces nothing No
NO_SIGNATURE No issue signature was supplied, so no correlation was possible No

Two consequences worth stating plainly:

  • A patch is never authored for a run that cannot prove its premise. The authorPatch callback is unreachable until the classification is REPRODUCED, so a run that proves nothing also spends nothing.
  • The classification travels onto the certificate, and auditPayload refuses any tiered claim whose qualification is not REPRODUCED — even when the MAC is valid. An operator can tell "the bug reproduced" from "the harness was broken" without reading a log.

A project with no reproCommandTemplate inherits node --test {file}. If its platform is not one Node can execute, the run stops at repro_command_supported before a single sandbox command is issued, naming the mismatch — rather than producing a non-zero exit that used to read as a reproduced bug.

Verification tiers

Tier Means Billable
A The repro reproduced the failure, the patch fixes it, and the full suite is green Yes
B The repro reproduced the failure, the patch fixes it, and the affected tests are green Yes
C The repro reproduced the failure and the patch fixes it; no suite was run No
NONE No verification claim is made No

Three rules that matter more than the table:

  1. The tier is what was ACHIEVED, never what was attempted. A run aiming for A that could not execute the suite reports B or C, and says why.
  2. Degradation is monotone downward. Tiers can only get weaker as a run proceeds; nothing can promote a result after the fact.
  3. A ceiling is computed from the suite configuration before the run starts. "This project can reach at most C, because no test command is configured" is a useful thing to say up front; "tier A" when the suite never ran is not.

Billing is on Tier A/B verification plus human merge only, with the artifact bundle attached to the invoice line — deliberately avoiding the "assumed resolution" billing disputes that dogged Intercom Fin.

Provenance certificates

Every proposal carries a signed, machine-readable attestation chain:

issue/event → repro test hash → patch hash → suite result → tier → approver
  • Self-contained. Everything needed to judge the claim is in the payload. A verifier needs the published key, not our database or our uptime. A certificate you have to call us to believe is a lookup, not evidence.
  • Canonical. Signing covers a deterministic serialization with sorted keys, so re-serializing in another language yields the same bytes and the same signature.
  • Tamper-evident in both directions. Verification checks the signature and the payload's internal consistency: a tier-A claim carrying no suite result is rejected even when the signature is valid. The interesting attack is a caller who signs an over-claiming payload, not one who edits bytes afterwards.
  • It never attests an environment that did not exist. environment.commit is only written alongside commitCheckedOut: true, which the loop sets from what the sandbox really did rather than from what the run asked for: the clone and the checkout both fail the run outright, so a signed certificate naming a revision is one the sandbox reached. A run whose provider cannot clone is not refused — it verifies the working tree it was given, says so in the run's degradations, and its certificate names no commit at all. Earlier certificates copied the requested commit through even on a provider that cannot clone, so they attested a checkout that never happened; a payload naming a commit without the flag now fails the audit with COMMIT_NOT_CHECKED_OUT.
  • It records why the pre-patch failure counted. repro.qualification carries the outcome, a stable reason code and one sentence of explanation.

GET …/proposals/:id/certificate returns it.

Rejection memory

"Never re-propose a rejected fix without new evidence" is the difference between an agent that learns and one that nags. Prose in a review comment cannot be enforced, so a rejection is compiled into machine-checkable constraints that the next run must load, cite and satisfy:

Constraint Forbids
exact-diff:<sha256> This exact patch, ever again
avoid-path:<glob> Modifying files matching the glob
forbid-approach:<slug> A named technique the reviewer rejected
require-tests A fix here without tests
needs-evidence:<issue> Re-proposing over the same files without new evidence

The vocabulary is deliberately small and readable: an operator looking at avoid-path:src/legacy/** in a list should know what it forbids without a manual.

The sandbox

Execution goes through a SandboxProvider seam. LocalSandboxProvider is the reference implementation; hosted providers are stubbed behind the same interface. The sandbox has an egress guard, because a patch generated from an error message is exactly the code path an injected instruction would try to exfiltrate through.

Sandbox verification is cloud-only. Community Edition ships no sandbox provider and refuses to start a run, with an explanation — see self-host-ce.md.

Approval

Merging is human-only and unconditionally so. pr.merge and fix.approve are denied to every non-human principal before any policy row is consulted — no configuration grants them. That is what makes "a human approved this fix" a fact rather than compliance theater. See agent-governance.md.

What is not built

Four entries here previously described the loop as inert — no deployed worker, no real patch author, no way for a diff to leave the database, no witness and no offline verifier. All four have since shipped and the list said otherwise, which is the exact failure this page exists to prevent. Corrected, with what is actually missing:

  • The hosted sandbox has no credentials here. LocalSandboxProvider really runs — it spawns the install command, the repro and the suite in a temporary workspace and tears it down — and the Modal adapter is a complete implementation rather than a stub. What this deployment does not have is MODAL_TOKEN_ID/MODAL_TOKEN_SECRET/MODAL_APP_NAME, so it refuses the first run by name instead of falling back to the local provider: the certificate records which sandbox verified a fix, and a quiet fallback would make that record false. Vercel stays a stub by decision — its sandboxes have no snapshots, so every repro would start cold. The local provider is refused in production unless SANDBOX_ALLOW_LOCAL=true, which does not configure a sandbox; it removes the only thing stopping tenant-supplied command strings from running on the worker host.
  • Certificates issued before 0.14.0 are MAC-signed and cannot be checked by anyone else. signCertificate now signs Ed25519 with the same key published at /.well-known/proof-ledger-keys, so a stranger holding only the public half can check a certificate. Dispatch is on the algorithm the certificate itself records, so older rows keep verifying with the secret that signed them, forever — and equally, remain forgeable by anyone who can verify them. Ask for a re-issue if you were handed one. The Proof Ledger bundle is the other artifact and is unchanged: Ed25519 over DSSE, checked offline by tooling/bucker-verify, which makes no network calls and can be pointed at a key you supply rather than the one embedded in the document.
  • No public transparency log is operated. Witness transports and co-signature verification are implemented, and a witness that fails to countersign is a hard delivery failure — but Bucker does not run a public log for anyone to watch, and the object-store witness currently writes to a filesystem blob store rather than an append-only object store.
  • Tier C is the default outcome. A project with no SuiteConfig row has a ceiling of C, and C is not billable — so a default deployment produces zero billable verified fixes until a suite is configured.
  • createPullRequest is not on any route. Branch creation, commit and pull-request opening are implemented against the GitHub Git Data API and are driven from the approved-proposal path; there is deliberately no endpoint that opens a PR directly. Providers other than GitHub return ScmNotImplementedError rather than pretending.
  • No auto-merge, at any tier, under any configuration. Deliberate — and structural: ScmClient has no merge method to call.
  • Cross-tenant fix memory (learning from other tenants' fixes) is deferred pending a consent model, in every edition.
  • No public per-customer fix success-rate dashboard yet.