← technical essays
[ESSAY]
No. 6.51 Mar 31, 2026 pillar essay

Observability Is the Second Product

Every running system is two products: the one users see and the one operators do.

[ essay ]

Thesis

Observability is a product with users: operators, on-call engineers, and future-you. Every system in production is two products. The first is what the reader sees. The second is what the operator can ask of the running machine. Shipping features without shipping the ability to ask novel questions is shipping half a product. The second half decides whether the first survives its second year.

Context

The mystic-bytes cover pipeline served stale thumbnails after a manifest correction. Users saw old art. The database had new paths. The CDN had an older cache generation. Public URLs had to stay stable. Debugging meant SSH, curl, a Redis CLI, and an argument about which layer was telling the truth.

We had logs. They were unstructured, unaggregated, and silent on success. We had no metric that said manifest version N is being served while the edge cache is still at M. The user-facing product looked fine if you did not refresh hard. The operator product was a rumor.

The fix for staleness was a contract rewrite. The fix for recurrence was observability: manifest version stamped on preprocess output, cache key logged on hit and miss, one dashboard panel asking whether those numbers were equal right now. No new vendor. Fields we should have had when we shipped the first product.

I have done this the other way. Shipped the pretty path. Promised myself I would “add metrics later.” Later arrived at 2am as a hypothesis I could not test without deploying a printf.

Mechanism

Charity Majors distinguishes monitoring from observability. Monitoring is someone else’s dashboards for failures you already named. Observability is the ability to interrogate production with a question you did not know to ask until the unknown unknown arrived.1 The Google SRE book codified the earlier layer: metrics, logs, and traces as the minimum operator surface.2 Both matter. The second product fails when teams stop at an uptime graph and call it done.

The second product has a feature list, whether you wrote it down or not. High-cardinality context: request id, slug, manifest version, cache tier — not only error_rate. Traceable paths: follow one thumbnail from import through preprocess to the edge. Actionable alerts: page when user-visible truth diverges from origin, not when CPU hiccups. Runbook links: every alert knows which doc teaches the next move. If those features are missing, you still have a second product. It is just a bad one, assembled in the incident channel, which is the worst IDE I know.

Reverse-engineering under fire is the failure mode. Teams that build only the user-facing product discover they cannot tell whether latency is auth, the database, or an N+1 introduced last Tuesday. They add logging while users wait. They learn the shape of the system from the outage, which is an expensive school.

On stream infra I treat dropped chat events like cover staleness. Instrument the relay queue depth and the plugin heartbeat in the same place I instrument viewer count. Viewers are the first product. The operator overlay is the second. When the only alert is “stream offline,” you learn about backpressure from angry chat, not from a graph.

Observability is an architectural constraint. If you cannot trace it, you do not understand it, and you should not pretend the merge is safe. Forcing every mystic-bytes merge job to emit a span with slug, batch_id, and cover_hash caught duplicate slug imports before they hit the manifest. We did not predict that bug. The question became askable, so the bug had somewhere to stand.

Logs are not free telemetry. Unstructured printf in a hot path becomes expensive noise you cannot query. Structured logs with stable field names — manifest_version, cache_tier, slug — let you pivot from a metric to evidence without a redeploy. The second product needs schema discipline the same way an API does. A log line that cannot be grouped is a diary entry. Diary entries do not page you with the right slug.

Alerts are product UX for a tired human. The user of an alert is someone who was asleep. The copy should name the divergence, the likely layer, and the first check. “CPU high” is a weather report. “Edge cache generation behind origin manifest for slug X” is a task. I would rather write one of those than ten charts nobody opens except to prove we have a dashboard.

Tradeoffs

Build vs buy. Hosted APM is fast to plug in. Cardinality limits and invoice shock show up when you finally label by slug. OpenTelemetry plus one backend is engineering tax up front and cheaper questions later. For mystic-bytes I stayed on what we already ran and spent the effort on fields, not on a new contract.

Cardinality vs cost. High-cardinality labels are how you debug. They are also how you bankrupt a metrics bill. Sample the boring success path if you must. Never sample errors. Never sample the one field that would have told you which cache tier was lying.

Noise vs silence. Alert fatigue kills the second product faster than a missing alert. One page you trust beats ten you mute. If the runbook for an alert is “look at the dashboard,” the alert is unfinished.

When thin observability is enough. Early prototypes and internal tools with one operator who holds the whole graph in their head can ship a feature plus a health check. The stance changes the day someone else carries the pager, or the day you will not remember the graph. That day arrives sooner than the roadmap admits.

Close

Ship the second product in the same stretch of work as the first. Ask what question you will need at 2am that you cannot answer today. Add the metric, the log field, or the trace span that makes it askable. Then ship the thumbnail.

If your on-call runbook says “check the dashboard” and the dashboard only charts user traffic, you have built half a product. The other half is what keeps the first one honest.

— JV · Dark Heart Labs.

References

  1. Charity Majors, George Miranda, and Liz Fong-Jones, Observability Engineering (O’Reilly, 2022). Majors’s distinction between monitoring known failures and asking novel questions of production is the reference for treating observability as its own product. ↩

  2. Betsy Beyer, Chris Jones, Jennifer Petoff, and Niall Richard Murphy, Site Reliability Engineering (O’Reilly, 2016). The Google SRE book is the canonical foundation for metrics, logs, traces, and SLO-driven operations as first-class work, not as a leftover. ↩

№ 6.51 — JV · Dark Heart Labs.