Skip to content
bucker

Rendered from docs/soc2-readiness.md in the repository, so this page and the assessment cannot disagree. Back to trust.

SOC 2 readiness

Status: not certified. No audit is in progress. No auditor has been engaged. There is no Type I report and no Type II report, and no observation window has begun.

That paragraph is the most important one in this document, and it is first so that nobody has to read to the end to find it. Everything below is an honest inventory: which SOC 2 Trust Services Criteria are already met by mechanisms that exist in this repository, which are met by intention rather than by tooling, and which have not been started. It is written for the person doing vendor diligence on a pre-launch company, who is better served by a map of the gaps than by a badge.

The convention throughout: a control is listed under In code only if a named file implements it and a named test exercises it. Anything else is under Policy, not yet tooling or Not started, even where the intent is settled.

Criteria references are to the AICPA 2017 Trust Services Criteria (2022 points of focus revision). Bucker has not selected a scope, so nothing here should be read as a commitment to which criteria a future engagement would cover.

The short version

Area Where it stands
Logical access (CC6) Substantially in code, and unusually strong: relationship-based authorization, human-only approval gates, per-agent blast-radius tokens
Change management (CC8) Substantially in code: no machine can merge, approvals are risk-tiered, evidence is signed and Merkle-logged
System monitoring (CC7) Partly in code: comprehensive audit trail, behavioural anomaly response; alerting on control failure is not defined
Confidentiality (C1) Partly in code: retention and deletion are implemented, but only the event stream is on an automatic schedule
Privacy (P) Partly in code: GDPR access/portability/erasure run against the live schema graph, and reach platform accounts only — the end-user data Bucker processes for a customer is served by a separate controller-scoped path, and object storage was outside every deletion path until 2026-09-09
Availability (A1) Not started, and deliberately: see status-page.md
Control environment, risk assessment, communication (CC1–CC5) Policy, not yet tooling — these are people-and-process criteria and Bucker has no security policy set, no risk register under maintenance, and no training programme

In code today

CC6.1 — Logical access is restricted by relationship, not by token alone

Authorization is ReBAC-shaped. A token's scopes act as a ceiling; the decision comes from the principal's relationship to the object — organization role, team membership, project grant — resolved from the database on every check.

  • code/domain/src/modules/authz/check.ts — check() applies both gates; a full-scope token with no membership is denied everywhere.
  • code/api/test/authz/check.test.ts — describe('scopes are a ceiling, never the decision').

Coarser standalone guards (requireScope, requireOrgRole) exist for routes with no object to relate to, such as organization settings; every path that reaches a project, issue or team object goes through the relationship check.

CC6.1 — Federated identity and directory-driven provisioning

SAML and OIDC federation, plus SCIM 2.0 directory sync covering /Users, /Groups and /Agents, on every paid tier rather than as an enterprise upsell.

  • code/domain/src/modules/sso/, code/domain/src/modules/scim/
  • Published documentation: code/docs/sso-and-scim.md

Deprovisioning is transactional: suspending a sponsor sets the agent's lifecycle, revokes its tokens, and revokes its delegations inside one database transaction (code/domain/src/modules/agents/identity.ts, suspendAgentsForSponsor), so an offboarded employee cannot leave a live machine identity behind.

CC6.1 — Delegation is bounded and cannot re-widen

Token exchange follows RFC 8693: nested actor claims, audience binding, a fifteen-minute default lifetime capped at one hour, and the presented token as its own ceiling — a narrowed session cannot mint a broader delegation than it holds (code/domain/src/modules/agents/delegation.ts; code/api/test/agents/delegation.test.ts).

Incident-scoped "blast radius" tokens bind an agent to one incident and its objects, are Ed25519-signed, and are re-verified by their own handler rather than trusting an already-attached principal (code/api/src/modules/agents/blast-radius.ts, code/domain/src/modules/agents/tokens.ts).

CC6.6 — Transport to the production database

The API dials the database over TLS with client-certificate material supplied as PEM text from encrypted secrets and passed to node-postgres as strings; nothing is written to disk, and there is no code path that reads a certificate from a file.

A DATABASE_URL naming a host reachable off the machine, with no TLS material configured, throws instead of connecting in the clear. Loopback, RFC 1918 and CGNAT addresses, .local/.internal-style names and single-label Compose service names are exempt, since those are not exposed transports; the escape hatch for a transport secured elsewhere is the explicit DATABASE_SSL_ALLOW_PLAINTEXT=true.

  • code/db/src/db-tls.ts, code/db/src/db-tls.test.ts
  • docs/mtls.md, docs/SECRETS.md §Mutual TLS

Precision worth stating: the code enforces encrypted and never on disk. Whether the connection is genuinely mutual is enforced by the database endpoint, which demands a client certificate — cert-and-key together is what our deployment configures, but the code would also accept server-authenticated TLS as "protected".

CC6.6 — Outbound requests are guarded against SSRF

Every outbound call from the API — webhooks, SIEM sinks, integrations — goes through a guarded fetch that refuses private, loopback and link-local destinations (code/shared/src/egress.ts).

CC6.7 — Personal data is removed before storage

The scrubber runs on the ingest path before anything is persisted, on all three entry points: the worker pipeline (code/workers/src/pipeline/process.ts), the agentless landing lane (code/domain/src/modules/landing/land.ts) and OTLP ingest (code/domain/src/modules/otlp/service.ts). There is no configuration flag anywhere that disables it: agent-observability's captureMode controls whether prompt and response text is retained at all, and even its full setting applies the stored-tier scrub (code/domain/src/modules/agent-observability/content-policy.ts).

A second, stricter pass runs before an issue reaches a model, at the "agent" tier, in the essence projection (code/domain/src/modules/essence/service.ts), in RCA evidence assembly (code/domain/src/modules/rca/evidence.ts), and in fix authoring (code/workers/src/remediation/execute-remediation-run.ts).

  • Rules and tiers: code/shared/src/scrub/rules.ts
  • Tests: code/shared/src/scrub/scrubber.test.ts
  • Replay-specific behaviour: code/docs/replay-privacy.md

A gap in the OTLP entry point, found on 2026-09-13 and closed. Naming OTLP above was only two-thirds true: otlp/service.ts scrubbed the tail that turns an exception-carrying log record into a grouped issue, and wrote Span.attributes, Span.resource, Span.name and every non-exception LogRecord.body/attributes exactly as the exporter sent them. That is where the OTel and GenAI semantic conventions put the dangerous content — db.statement, http.request.header.authorization, a url.full carrying userinfo, gen_ai.prompt holding a whole prompt — so the one ingest path with no SDK-side masking in front of it was also the one with no stored-tier scrub behind it. The rows now go through the same stored tier and the same per-org rules (FR-102) as every other path, per ROW rather than per batch because MAX_SCRUB_NODES is a per-call ceiling and one call over a 5,000-span batch would truncate its own tail. Tests: code/api/test/otlp/attribute-scrub.test.ts, which asserts the stored row AND the read API, and which was watched failing against the unscrubbed path first.

A gap this note found, and the fix. Writing this section surfaced a real defect: the remediation fix-author path built its prompt from the denormalized issue row (title, exceptionType, exceptionValue, culprit) with no agent-tier re-scrub, while code/workers/src/remediation/author-llm.ts interpolates those columns verbatim into the prompt. The stored tier removes secrets and card numbers but keeps email addresses, IP addresses and user identifiers deliberately, so an email in an exception message reached the fix-authoring model on every Tier-A run.

It is fixed: the context is now scrubbed at the agent tier, salted per project, in code/workers/src/remediation/execute-remediation-run.ts. The regression test in code/workers/test/remediation/orchestrator.test.ts was verified to fail when the tier is weakened back to stored, and it also asserts the non-PII half of the message survives, so an over-broad scrub that leaves the author nothing is caught too.

The history is left here rather than replaced with the fixed state, because the fact that writing this note found the defect is the argument for keeping the note honest.

CC6.8 — Secrets are encrypted at rest, and the guard is mechanical

There is no plaintext .env in this repository and no command that writes one. Secrets live in secrets/dev.enc.env and secrets/prod.enc.env under sops + age and are decrypted straight into a process's environment.

A pre-commit hook blocks secret-shaped paths, private-key material, and exact live secret values matched by SHA-256 against a local fingerprint index. It never decrypts anything, so committing does not prompt and the guard can run in contexts that hold no key.

  • scripts/secrets.sh, .githooks/pre-commit, .githooks/_guard.sh
  • scripts/test-secret-guard.sh — the guard's own test battery
  • Rationale: docs/SECRETS.md, docs/encrypted-secrets-recipe.md

CC7.2 — Anomalous machine behaviour is detected and acted on

Behaviour baselining learns an agent's normal shape from its own audit trail, excluding the anomaly window so an attack cannot teach the baseline, and revokes tokens on a high-severity anomaly, suspending the agent outright on two independent ones (code/domain/src/modules/threat/).

Two kill switches gate every agent action ahead of the relationship lookup: organization-wide AI enablement, and agent pull-request creation separately (code/domain/src/modules/authz/agent-policy.ts, agentDenialReason).

CC7.3 — The audit trail is action-granular and has one write path

Every audit row in the system is written by a single function, writeAuditLog (code/domain/src/modules/authz/audit.ts) — it is the only caller of prisma.auditLog.create anywhere in the API. Roughly 185 call sites across the modules reach it, directly or through four thin per-module wrappers, recording what was done, by whom, on whose behalf, and under what basis.

The single chokepoint is the control worth citing, not the count: it means audit coverage is reviewable by reading one function's callers, and a module cannot invent its own unlogged write path.

The Human Accountability Ledger reads that trail back per principal, with approvals and declines weighted equally and each decline classified by why (code/domain/src/modules/accountability/). A plain-language sentence renderer turns one event into one readable line (code/shared/src/audit-sentence.ts).

CC7.4 / CC7.5 — Public disclosure of our own miss rate

An unauthenticated, deployment-wide transparency register publishes rejected-patch rates, in-window regression rates and prompt-injection trial results, Merkle-logged with signed snapshots, inclusion proofs and consistency proofs, and with no edit, delete or retract route for a published snapshot (code/domain/src/modules/transparency/).

This is not a SOC 2 requirement. It is included because it is the strongest available evidence for CC4-style monitoring: the numbers that would embarrass us are the ones we publish without asking a reviewer to sign up first.

CC8.1 — No machine merges, and the change record is signed

Approving a fix and merging a pull request are refused to non-human principals. Three independent layers hold that line:

  1. The approval endpoint checks principal.kind !== 'HUMAN' first, with no database read before it (code/domain/src/modules/approvals/requests.ts, assertMayDecide). The shared decision engine applies the same gate, and its outcome never depends on the policy row (code/domain/src/modules/approvals/engine.ts).
  2. The fix:approve scope is filtered out of every agent token at issuance, and the pr:merge tier is ungrantable to an agent (code/shared/src/scopes.ts, code/domain/src/modules/agents/identity-tier-scope-vocabulary.ts).
  3. There is no merge method in the source-control client to call (code/domain/src/modules/scm/provider.ts, github-client.ts).

Any future action named merge or ending in .merge is treated as human-only by default (engine.ts, isHumanOnlyAction) — the fail-closed direction.

Approval gates are risk-tiered by construction rather than configurable to zero, and the approvals engine records approverInDelegationChain on every vote, so a blocked self-approval is a distinguishable outcome in the ledger rather than a silent pending row.

CC8.1 / PI1.4 — Change evidence is signed and logged

A verified fix emits a provenance certificate carrying its verification tier, and the evidence is published to an append-only Merkle log (RFC 6962 leaf and node domain separation) with Ed25519-signed checkpoints, inclusion proofs and consistency proofs (code/domain/src/modules/proof-ledger/, code/shared/src/proof-ledger/).

Two honest qualifications a reviewer should hold onto:

  • Only the Proof Ledger bundle and the checkpoints are third-party verifiable. They are Ed25519 over DSSE. The per-fix provenance certificate and the accountability-ledger export are HMAC-SHA256 keyed by the deployment's own secret: anyone who can verify one can forge one. They are internal integrity checks. The in-browser verifier refuses a pasted provenance certificate by name for exactly this reason (code/frontend/src/lib/proof-ledger/verify-in-browser.ts).
  • Witnesses are opt-in and default to none. Independent co-signature is implemented, with real Ed25519 countersignature verification against a registered key, receipts and retries (code/domain/src/modules/proof-ledger/witness.ts). But a customer must configure a witness, and a witness under Bucker's own custody is reported as not independent. The API exposes a headline independentlyWitnessed boolean rather than letting anyone infer independence from the presence of rows.

Redaction preserves the tree: redacting an entry nulls the readable envelope and stamps a reason, and leaves the leaf hash, every published root and every previously issued inclusion proof verifying unchanged (code/domain/src/modules/proof-ledger/issue.ts; asserted in code/api/test/proof-ledger/routes.test.ts). Entries containing personal data are refused at issuance rather than silently edited afterwards (code/domain/src/modules/proof-ledger/personal-data.ts).

C1.1 / C1.2 — Retention and disposal

Event data is retained per policy and actually deleted: the retention pass archives a day-partition to cold storage and then DROP TABLEs it, and expires cold objects past their window. It runs hourly by default and drops by default (code/domain/src/modules/retention/pass.ts; code/workers/src/retention/scheduler.ts).

Gap. The compliance retention module — audit logs, spans, log records, replay segments — implements correct, tested hard deletion, but its sweep has no scheduler. Today it runs only when an administrator calls POST /orgs/:orgSlug/compliance/retention/enforce (code/domain/src/modules/compliance/retention.ts). A retention policy with an operator-triggered sweep is not the same control as an enforced window, and a reviewer should treat it as the weaker of the two.

P — Data-subject requests

Access, portability and erasure are implemented and organization-scoped, with an admin authorization check before any data is touched (code/domain/src/modules/compliance/dsr.ts, routes.ts).

The subject surface is derived from the live foreign-key graph rather than a hand-maintained list of tables, so a new table referencing a user is covered the day it is added; the test asserts completeness against that graph (code/api/test/compliance/dsr.test.ts).

What that surface does not reach, stated plainly because the omission is the part a reviewer will ask about. Deriving the subject from the FK graph rooted at User/Principal covers Bucker's own platform account holders — the customer's staff who sign in. It does not cover the personal data Bucker processes on the customer's behalf: Event payloads, Span and LogRecord attributes, ReplaySegment recordings, IssueEmbedding vectors and EventArchive blobs all key on projectId only, hold no foreign key to a user, and are never scanned for a subject identifier. Compounding it, code/shared/src/scrub/rules.ts deliberately does not scrub email or IP at the stored tier — a documented choice, made for on-call debugging — so real end-user identifiers sit in the clear in exactly the tables this pipeline cannot see.

So a request from an end user named inside a customer's error payload is served by the separate controller-scoped erasure path, not by this one, and rows already delivered to a customer's own warehouse sink are outside Bucker's control entirely. An erasure that silently misses data is worse than one that declares its boundary, which is why the boundary is enumerated in the response rather than implied.

Erasure is a policy, not a blanket delete, and the shape matters to a reviewer:

  • Nullable references to the subject are set to null.
  • A leaf row whose reference is non-nullable is deleted.
  • A non-leaf row whose reference is non-nullable keeps its pointer, aimed at a pseudonymized tombstone.
  • The user row itself is pseudonymized rather than deleted.
  • Audit logs are retained, under GDPR Art. 17(3)(b)/(e), reduced to a principal id and a SHA-256 input digest — never raw payload.

Sub-processor and data-flow transparency

Audit events can be streamed to a customer's own SIEM over an HMAC-signed NDJSON webhook, with a durable cursor that advances only on successful delivery (code/domain/src/modules/compliance/audit-export.ts). One caveat: the sink kind is a generic WEBHOOK — there is no Splunk, Datadog or Elastic-specific driver — and delivery is flushed on demand rather than on a scheduler, so this is not yet a continuous stream.

Policy, not yet tooling

These are intentions with owners and no implementation. Listing them as controls would be the exact failure mode this document exists to avoid.

  • CC1 (control environment) — no board or independent oversight function; no code of conduct; no background-check process; no defined security roles beyond the repository's ownership conventions.
  • CC2 (communication) — no published internal security policy set; no security awareness training; no formal channel for reporting control failures. The engineering-facing documentation in code/docs/ and docs/ is genuinely maintained and is the closest thing that exists.
  • CC3 (risk assessment) — the delivery plan carries a risk register (docs/2026-08-28-plan.md §13) covering engineering-delivery risk. It is not a security risk assessment, it has no review cadence, and no fraud-risk consideration exists.
  • CC5 (control activities) — control design is embedded in code review and the ownership manifests rather than documented as a control matrix.
  • CC9 / vendor management — no maintained sub-processor inventory, no vendor security review process, no data-processing agreements.
  • Incident response — no written IR plan, no severity definitions, no on-call rotation, no customer notification commitment or timeline.
  • BCDR — there are no backups. This line previously read "backups exist as a property of the managed database, untested by us", which was the one substantive overclaim in this document: nobody has scheduled a dump, and code/deploy/fly/README.md's own "Still not set up" section agrees. A provider's snapshots are not a tested restore, and this deployment has neither. There is also no documented recovery objective and no restore drill. Compounding it, the secret store has one age recipient and no escrow, so a backup nobody can decrypt would not be one either.
  • Access review — SCIM makes deprovisioning mechanical, but there is no periodic entitlement review of who holds which organization role.
  • Change management, human side — the technical gates are strong (above); there is no documented change advisory process, no defined emergency-change path, and no separation of duties statement covering who may deploy.

Not started

  • Selecting an auditor, or a readiness assessment with one.
  • Scoping: which criteria, which systems, which period.
  • A Type I point-in-time examination.
  • An observation window, and therefore any Type II report.
  • Evidence-collection automation (a compliance platform, or scripted evidence export).
  • Penetration test by a third party.
  • A published sub-processor list and a customer-facing security page beyond /trust.
  • Availability commitments of any kind: no SLA, no SLO, no uptime history. See status-page.md for the deliberate decision not to build one.

What would have to happen next

In rough order, and stated so that a reviewer can judge the distance rather than take a timeline on faith:

  1. Close the two gaps named above — agent-tier scrub on the fix-author prompt, and a scheduler for compliance-stream retention. Both are engineering work already scoped by the code that surrounds them.
  2. Write the policy set (CC1, CC2, CC5) and the incident-response plan. This is the bulk of the remaining work and none of it is code.
  3. Stand up a maintained security risk assessment and an access-review cadence.
  4. Engage an auditor for a readiness assessment; scope the criteria.
  5. Start an observation window. A Type II report cannot exist earlier than the end of it, and no amount of prior engineering shortens it.

Provenance of this document

Every claim in the In code section was checked against the repository, not against the roadmap, before it was written; the file paths are given so the check is repeatable. Where a claim in an earlier draft did not survive that check it was weakened or removed rather than softened — the fix-author scrub gap, the retention scheduler gap, the witness default, and the HMAC/Ed25519 distinction are all here because of it.

This document has no auditor's opinion behind it. It is an engineering self-assessment, and it should be read as one.