Skip to content
bucker

Session replay: what is captured, what is masked, and where scrubbing happens

Audit of the replay privacy pipeline, SDK capture through stored segment. The code of record: sdks/browser/src/replay/ (capture and client-side masking), sdks/browser/src/microdom.ts (selector construction), domain/src/modules/replay/ingest.ts (server-side scrub and storage), domain/src/modules/replay/transcript.ts + domain/src/modules/essence/trust.ts (agent-facing compilation and injection quarantine).

The recorder is not rrweb. It records the interaction layer — no DOM mutations, no screenshots, no text nodes except where explicitly noted below. Most of the privacy story is therefore structural: the bytes for a pixel-accurate replay of the user's screen are never collected in the first place.

1. What the recorder captures, per event type

Event Captured User content?
nav route template (/orders/:id), previous template No — ids collapsed, query string dropped
click micro-DOM selector chain (form#pay > button#submit) Element label text only if maskAllText: false
input selector, masked/sensitive flags, value Value is <masked> unless maskAllText: false
net method, URL template, status, duration, error name No — query strings dropped, path ids collapsed
console console.error/console.warn arguments, stringified, ≤200 chars Yes — captured verbatim at every setting (see §4)
scroll coalesced scrollY No
resize viewport width × height No

Nothing leaves the page until an error is captured: the ring buffer (buffer.ts) holds at most the last 30 seconds / 200 events in memory, and an error-free session is never transmitted. Eviction is counted (evicted), so a short segment is explainable rather than suspicious.

Client-side masking (mask.ts)

  • Default is mask-everything. maskAllText defaults to true; input values arrive as the literal <masked> and click labels are omitted.
  • Credential-shaped fields are masked at every setting. isSensitiveField (type=password, autocomplete credential tokens, cc-*, one-time-code, name/id/aria-label matching password/token/cvv/ssn/… shapes, data-bucker-mask / data-private / data-sentry-mask attributes) is not reachable from any option, by construction — turning masking off is not consent to ship a password.
  • URLs are templated before capture. urlTemplate drops every query string and collapses numeric / UUID / long-hex / opaque path segments to :id, so a /reset-password/<token> never leaves the page as a live credential. Cross-origin hosts are kept (which third party failed is load-bearing).
  • Selectors are attribute skeletons. microdom.ts keeps tag, id, data-testid, role, name, type and one build-stable class; hashed class names are pruned. Element text is read into the micro-DOM node but only ever emitted on click.text, which is gated on maskAllText: false.

2. What reaches the server

One bulk-lane envelope per error (rate-limited to one segment per 5 s), carrying:

  • replay_event — segment metadata: replay/segment ids, time bounds, event count, starting URL template, masked flag, eviction count, stop reason, overhead.
  • replay_recording — the event array itself.

Both are re-validated on the server against schemas declared server-side (replay/schemas.ts): a replay payload is authored by a script running on the customer's page and is attacker-controllable in the strongest sense. Strings are bounded, events capped at 2000/segment, and an invalid item stores nothing.

3. Where scrubbing happens relative to storage

page ──mask (SDK, capture time)──▶ wire ──▶ accept (counted, not parsed)
  ──▶ worker ingest:  zod validation
                      ├─ SCRUB, stored tier  ──▶ blob write (payloadKey)
                      │                        ──▶ ReplaySegment row (url, meta)
                      └─ SCRUB, agent tier   ──▶ transcript compile
                                                  ├─ injection quarantine (Labeler)
                                                  └─ ReplayTranscript row
  • Stored tier before any byte persists (ingest.ts): the same rule set that guards the Event row — JWTs, Bearer/Basic values, vendor tokens (sk-…, ghp_…, xox…), AWS key ids, private-key blocks, card numbers (Luhn-checked), high-entropy secrets — runs over the event array and the metadata URL before the blob write and before the segment row. The redaction summary (e.g. 2x jwt) is stored inside the blob and served with the payload, so a [Filtered] value is explainable to whoever is looking at it.
  • Agent tier before the transcript: the transcript is agent context by definition, so it compiles from the strictly narrower agent projection — additionally redacting emails and public IPs that a human on-call engineer may legitimately see in the payload view. The same applies to …/replay/anchors labels, which ride RCA evidence: they are built from agent-tier-scrubbed events at read time.
  • Injection quarantine at compile and at anchor build: every page-derived string served to an agent goes through the essence Labeler; matches are replaced with [redacted:quarantined] and the document publishes a securityNotice. Quarantine is about instructions, the scrub about data — they are independent and both mandatory.

Storage split: the recording lives in the blob store under replays/<projectId>/<day>/… (retention = drop the day prefix); Postgres holds only metadata (ReplaySegment) and the compiled, already-redacted transcript (ReplayTranscript). Segments are bounded (≤2000 events, every string bounded), so a payload tops out at low single-digit MB and is served whole, JSON, sorted by the recorder clock — no chunking API is needed at these sizes.

4. Findings

F1 — fixed server-side: replay bypassed the pipeline's scrub stage. Error events pass through scrubEvent before storage (workers/src/pipeline/process.ts), but replay envelopes went from zod validation straight to the blob store: a Bearer token in a console.error, or a secret pasted into an unmasked input, was stored verbatim and could surface in the transcript (the injection scanner does not catch credentials). Fixed in ingest.ts: stored-tier scrub before the blob write, agent-tier scrub before transcript compilation, per §3.

F2 — SDK-side, flagged (out of scope to change here): maskAllText does not apply to console capture. recorder.ts records console.error/console.warn arguments verbatim (≤200 chars) regardless of the masking setting, while the segment's masked: true flag suggests otherwise. Applications routinely interpolate emails, names and order details into error logs, so a "masked" segment can still carry user text in console.message. The server-side scrub now removes credential-shaped content, and emails never reach the agent projection, but stored blobs retain what the page itself logged. Recommended SDK change: honor maskAllText for console messages (or add a dedicated maskConsole option) and reflect the residual exposure in the masked flag.

F3 — SDK-side, minor: selector attributes are captured unmasked. id, name and data-testid values ride into selectors and are not maskable. In practice these are developer-authored, but an app that writes user data into DOM ids would leak it into selectors. The server bounds them (≤512 chars) and scrubs credential shapes; no further action recommended beyond documenting the assumption.

F4 — accepted risk, mitigated: the metadata URL is client-asserted. Our SDK sends a template, but the schema cannot force a hostile or legacy client to template. Since the scrub now runs over meta.url too, a raw /reset/<jwt>?email=… is redacted before it reaches the ReplaySegment.url column; a raw email in a path segment, however, would persist at the stored tier (consistent with how event URLs are treated).

F5 — honest-flag note: masked is asserted by the client. The row's masked column reflects what the SDK claims, not what the server verified. It is surfaced as an honesty field for triage, not as a guarantee; the guarantees are the ones enforced server-side in §3.

5. What the read API serves

  • GET …/issues/:id/replay — segment metadata only (time range, duration, event count, size, masked, evicted, transcript availability). No payload bytes.
  • GET …/issues/:id/replay/transcript — the agent form: agent-tier scrubbed, quarantined, token-budgeted, provenance-labeled per line.
  • GET …/issues/:id/replay/anchors — notable moments (error, navigation, last-N interaction) as { timestampMs, offsetMs, label }; labels are built from agent-tier-scrubbed events and labeled/quarantined like transcript lines.
  • GET …/replay/:segmentId — the stored-tier payload (humans may see emails here; agents are never handed this route's output directly), served as JSON with its redactions summary and format descriptor (bucker.replay.interaction + version) so a player can refuse formats it does not know.

All four routes authorize as read on the issue, 403-not-404 for permission denials, with existence resolved tenant-scoped first (a foreign segment id 404s before any authorization question is asked about it).