Observability — the map
Three planes, one join key each, and no second vendor. This document is the map: what the platform emits, where it lands, which question it answers, how to get from one plane to another, and what is still blind. The event taxonomy for the product plane is its own document — analytics-event-model.md.The three planes
They answer different questions and are not interchangeable. A funnel cannot
tell you why a turn died; a trace cannot tell you that eleven workspaces hit the
same 500; a bug map cannot tell you whether revenue moved.
Why no vendor for the failure plane
The repo’s standing rule — stated ininvestigation-operations.ts and pinned by
its tests — is that a second observability vendor means a second definition of
“a failure”, and the day the two disagree is the day nobody trusts either. So
the failure plane is a table this platform writes plus SQL an operator runs. No
client, no transport, no agent, no per-seat bill, and nothing to keep in sync.
The failure plane
A bug is a fingerprint, not an occurrence
This is the idea the whole plane rests on. A log line per failure is a stream: readable, but impossible to count, rank or trend, and useless for “is this new?” or “how many customers does it hit?”. A fingerprint turns the same failure seen a thousand times into one thing with a count.- the message — it carries values (
Topic 'billing-flow' not found), so it would split one bug into one bug per value. It is also where customer data lives, and this plane holds none. - the release — a bug that survives a deploy must stay the same bug, or “did this deploy introduce it?” is unanswerable. Release is a column.
- the workspace — one bug hitting thirty workspaces is one bug with a blast radius of thirty, not thirty bugs. Workspace is a row dimension.
GET /workspaces/acme/topics/x and
GET /workspaces/other/topics/y were different endpoints — every workspace was
its own bug and every error breakdown was a list of one-row groups. See
ops/failure-fingerprint.ts.
The ledger is rolled up
error_events holds one row per (day, workspace, fingerprint) and increments
occurrences. That is the design, not an optimisation:
- it grows with the number of distinct bugs, not with traffic, so an error storm cannot become a storage incident;
- ten bugs is ten rows a day no matter how hard the platform is being hit;
first_seen_atis never updated, which is what makes new and regression answerable at all.
last_correlation_id away.
The writer cannot make an incident worse
ops/failure-ledger.ts sits on the path of every failed request, which means it
runs hardest exactly when the system is least healthy — and the database it
writes to can itself be the failure. Three defences, all load-bearing:
- Never awaited. A slow write cannot add latency to an error response.
- A circuit breaker. After 5 consecutive write failures it stops for 60s. Without it, a database outage means every 500 attempts another write to the database that is down — one outage amplifying into a write storm.
- A rate cap. 20 writes/second/process. The ledger is a map, not a tape, so it may shed — and what it sheds is counted and logged, never hidden.
ledger_health view: its counts are a
floor, not a total. A view that let them read as totals would be lying during
exactly the incidents that matter most.
The views
Inops/failure-operations.ts. SQL text plus typed row shapes — the audience is
an operator at a psql prompt or a BI tool. $1 is always the window start.
Unlike the investigation views, these are platform-wide, not workspace-scoped.
A bug belongs to the engineering team, and its most important property — blast
radius — is unanswerable from inside one workspace. The workspace appears as a
count(DISTINCT …), never as a filter.
failure_surfaces exists because the HTTP plane is only part of the picture: a
chat turn that dies is never a 500. It returns a polite sentence and a failed
row in agent_runs. A map that showed only error_events would report a healthy
platform while every turn was failing.
Getting from one plane to another
correlation_idis minted per request byCorrelationInterceptor, returned onX-Correlation-Id, printed on every log line, sent asrequest_idonrequest_failed, and stored aslast_correlation_idon the bug row. A customer can paste it into a support message and it resolves to the exact failure.run_idis theagent_runsprimary key. It is on the Latitude span (driftless.run_id), and now onchat_answer_received/chat_answer_failedtoo — without it a funnel drop was a number nobody could open.
Known blind spots
Stated because a map that hides its edges is worse than no map.- The dashboard and the marketing site have no failure plane. A client-side exception, a failed fetch, a chunk that will not load: none of it reaches any of the three planes. The API sees only requests that arrive. This is the largest remaining gap.
correlation_idis not on the Latitude span. Trace ↔ log joins throughrun_idfor agent runs; for a plain HTTP request there is no span at all.- Counts are a floor whenever the rate cap bites — by design, and
ledger_healthis how that is noticed. - 404 and 401 are not captured, deliberately: they are probe and
missing-resource noise that would drown the signal (
shouldCapture). A real 404 bug is invisible to the map. - No alerting. The plane makes bugs visible; nothing pages anyone. A query
on
new_bugsis the natural first alert, and it does not exist yet. - Background jobs and webhooks raise failures that never become an HTTP response — the cron reapers, the Stripe and broker webhook handlers. Their failures land in the log only.
- No retention job. The roll-up bounds growth by construction, and
RetentionJobis deliberately not extended: a candado (retention-scope.guard.spec.ts) makes a fourth purge target fail loudly, and the honest answer is that this table does not need one.
Adding to the map
- If it is a funnel question → the product plane, and the event goes in the
registry (
libs/analytics). - If it is a “what happened inside this run” question → a span attribute in
libs/telemetry. - If it is a “what is broken” question → a view in
failure-operations.tsover a ledger that already exists. If no ledger carries it, that is a real gap worth a table — but the bar is that no existing table can answer it.
