Staged-rollout guardian
A phased mobile release goes to 1–5% of users first. The guardian watches that cohort's crash-free rate and stops the rollout when the new build is worse than its predecessor was at the same point in its own rollout.
Comparing against a fully-adopted previous release flatters every new build, because the users who adopt first are systematically the stable ones. So the comparison is anchored on each release's own hour zero.
The decision
Two gates, both required before anything is called a regression. They pull in opposite directions on purpose: a guardian that misses a real regression is useless, and a guardian that halts on noise gets switched off — after which it is also useless.
- Sample floor — both cohorts need
minSessions(default 500). A 1% cohort with 20 sessions and two crashes reads as a 10-point crash-free drop. - Significance — a one-sided two-proportion z-test on the crash counts, with a p-value ceiling (default 0.01). A large cohort still produces sub-percent deltas by chance. The effect threshold (default 0.5 crash-free points) says "big enough to care"; the p-value says "not luck". Both must clear.
Verdicts: HEALTHY, INSUFFICIENT_DATA, NO_BASELINE, INCONCLUSIVE, REGRESSION.
Only REGRESSION can halt anything, and every non-halt records why.
The crash-free arithmetic is not reimplemented here — it is the same
release-health code the release dashboards read. The number on the screen and the
number the guardian halts on are the same number, computed once.
A halt needs a human
rollout.halt is unclassified in the risk table, so it falls to SENSITIVE → ASK
under the default approval policy. The guardian never pulls the cord by itself.
Stopping a release that turns out to be fine has a real cost, and a guardian that can halt unattended is one bad statistic away from an outage of its own making. The approvals engine also enforces its usual rule: the initiator of a request may not be its sole approver.
The flow is one idempotent endpoint:
| State | POST …/rollouts/:id/halt does |
|---|---|
No regression, no force |
Nothing. Returns the assessment and why it did not act. |
| Regression, no approval yet | Opens an approval request, sets the rollout HALT_PENDING. Nothing else happens. |
| Approval still pending | Returns the same request. No duplicate is opened. |
| Approval granted | Executes the halt. |
| Approval rejected | 409, and the rollout returns to ACTIVE. A new request must be opened deliberately. |
| Org policy says ALLOW | Executes immediately. |
What a halt produces
In this order, and the order is the durability contract — everything before the provider call is ours and cannot be lost by a receiver being down:
- A bisect toward the suspect release and commit (below).
- An issue, on the same
Issuetable as everything else, via the normal ingest path. One lifecycle, one triage surface, no parallel "rollout incident" model. The fingerprint is the rollout, so repeated halts collapse into one issue rather than a storm. - A ChangeEvent (
DEPLOY), because "the rollout of 4.2.0 was halted" is exactly what the RCA engine reads as evidence when explaining a later error onset. - The rollout marked
HALTED. - The provider called. A receiver that 500s is reported as an unconfirmed halt, never as a lost decision.
- An alert, through the one routing engine. The guardian has no channels of its own.
Bisecting
Which release. Releases are ordered by when they actually started receiving
sessions — never by version string, because 1.10.0 vs 1.9.0 vs 2.0.0-rc1 has no
ordering everyone agrees on. The oldest release in the probe window is the anchor,
and every later release is compared against it with the same guardian. The suspect is
the first release that regressed against the anchor.
Anchoring on the oldest rather than on each release's immediate predecessor is what makes the answer right when badness persists: if 4.1.0 broke it and 4.2.0 merely inherited the same crash, comparing 4.2.0 against 4.1.0 looks flat and finds nothing. 4.1.0-against-4.0.0 is where the drop actually is, and that is the diff someone needs.
This is a bounded linear scan, not a binary search, despite the name. Binary search needs a monotone predicate and crash-free rate is not monotone over release order — a project can ship bad, fixed, bad again.
Which commit. Inside the suspect release, the issue that first appeared in that window with the most affected users becomes the anchor, and the SCM module's existing suspect-commit ranking runs on it — frame depth 0.70, recency 0.20, file overlap 0.10, with a plain-English reason on every candidate. That ranking is not reimplemented; a second blame heuristic that disagreed with the issue page would be worse than none.
An empty answer is a valid answer. No prior releases, no code mappings, no in-app frames → no suspects, with a sentence saying which. Fabricated blame costs an agent its whole context window in the wrong file.
Providers
| Provider | State |
|---|---|
generic |
Works. Signed webhook: POSTs a rollout.halt signal to a URL you control; your CI/CD, flag service or fleet controller performs the pause. |
play |
Not implemented. Halting a Play track needs an Android Publisher service account and an edit transaction we have not built. |
appstore |
Not implemented. Pausing a phased release needs an App Store Connect API key and ES256 JWT minting we have not built. |
The stubs throw rather than quietly returning "not halted". A provider that appears
to halt a public rollout but does not is worse than one that refuses.
GET /rollout-providers reports implemented per provider so the gap is visible
before 2am, and creating a rollout on an unimplemented provider is refused at creation
time with the working alternative named.
Everything upstream of the store call — detection, approval, the issue, the ChangeEvent, the alert, the bisect — runs identically for a Play rollout today. Only the final "pause the track" call is missing.
The halt signal
POST https://ci.example.com/hooks/rollout
content-type: application/json
x-bucker-event: rollout.halt
x-bucker-signature: t=1786000000,v1=<hex hmac-sha256>
{
"version": "1",
"event": "rollout.halt",
"rolloutId": "rol_…",
"projectId": "prj_…",
"releaseVersion": "4.2.0",
"environment": "production",
"externalRef": "com.acme.storefront",
"cohortFraction": 0.05,
"reason": "4.2.0 is 94.00% crash-free at 6.0h vs 99.50% for 4.1.0 at the same stage …",
"suspectRelease": "4.2.0",
"requestedAt": "2026-08-07T16:00:00.000Z"
}
The signature is the same t=…,v1=… scheme as alert webhooks — the timestamp is inside
the MAC, so a valid body cannot be re-dated — which means a receiver that already
verifies Bucker alerts verifies this with the code it has.
API
GET /rollout-providers
GET /orgs/:org/projects/:project/rollouts
POST /orgs/:org/projects/:project/rollouts
GET /orgs/:org/projects/:project/rollouts/:id
GET /orgs/:org/projects/:project/rollouts/:id/checks
POST /orgs/:org/projects/:project/rollouts/:id/evaluate
POST /orgs/:org/projects/:project/rollouts/:id/halt
POST /orgs/:org/projects/:project/rollouts/:id/status
GET /orgs/:org/projects/:project/rollouts/:id/bisect
read to look, write to create or evaluate, deploy_hotfix to halt — halting is
the same authority class as pushing a hotfix, because both change what users are
running right now without the normal release process. The approvals engine then gates
it again with a human vote; authz decides who may even ask.
Creating a rollout freezes its baseline (the preceding release by observed rollout order). If the baseline were resolved live, a newer release starting mid-rollout would silently change what the guardian is comparing against.
Every probe is recorded, including healthy ones. "The guardian was watching and saw nothing" is precisely the fact an incident review needs, and a table containing only halts cannot answer it.
Not built
- No scheduler.
evaluateis called by you (cron, CI, the dashboard); there is no background sweep yet. - No automatic cohort-fraction tracking.
cohortFractionis what you told us. - Apple's phased release is day-quantised (1/2/5/10/20/50/100% over seven days); the guardian's equivalent-stage window is in hours and does not yet map onto phase days. That mapping is part of the missing App Store provider, not an afterthought.