Observability Is Not Logging: What to Instrument Before Something Breaks
Console logs and uptime pings don't tell you why a system failed. Here's what real observability requires, and why it pays for itself the first time you need it.
Most teams find out their observability is inadequate at the worst possible moment: production is down, a customer is asking what’s happening, and the only tool available is grep-ing through logs that were never structured to answer the question being asked. The team ships a fix, the incident closes, and everyone quietly agrees to “add better logging” — a promise that rarely survives contact with the next sprint.
The problem isn’t a lack of logs. Most systems produce plenty of them. The problem is that logging, monitoring, and observability are three different capabilities, and having one doesn’t give you the other two.
Logs, Metrics, and Traces Answer Different Questions
A log tells you that something happened: a request came in, a job failed, an exception was thrown. It’s a record of discrete events, and it’s the tool most teams reach for first because it’s the cheapest to add — one line in the code, no infrastructure decision required.
A metric tells you how something is trending: request latency over the last hour, error rate by endpoint, queue depth over time. Metrics are cheap to store and fast to query precisely because they throw away detail, which makes them the right tool for dashboards and alerting, and the wrong tool for figuring out why one specific request was slow.
A trace tells you where time went inside a single request as it crossed service boundaries — which downstream call, which database query, which retry loop actually caused the latency. This is the piece most teams skip, because it requires propagating a trace ID through every service a request touches, and that only works if it was designed in from the start. Bolting tracing onto a system after an incident means instrumenting under pressure, which is when it’s most likely to be done badly.
None of the three substitutes for the others. A system with rich logs and no metrics can tell you an error happened but not whether it’s getting worse. A system with metrics and no tracing can tell you p99 latency spiked but not which downstream dependency caused it. Observability is having all three wired together well enough that a question you didn’t anticipate can still be answered from data you already collected.
Instrument Before You Need To, Not After
The test of good instrumentation isn’t whether it exists — it’s whether it answers a question nobody thought to ask in advance. That’s only possible if the data was structured for querying, not just for reading. A log line that says "payment failed" is nearly useless during an incident; a structured log with event: payment_failed, user_id, provider, error_code, and latency_ms lets you filter, group, and correlate without redeploying anything.
This is why instrumentation has to be a design decision, not a cleanup task. Teams that treat it as a design decision add structured fields, trace propagation, and key metrics as part of building a feature — the same way they’d add input validation. Teams that treat it as cleanup add it after an outage exposes the gap, which means the next unanticipated failure mode exposes a new gap, indefinitely.
Alerts Are a Tax on Attention — Spend It Carefully
More monitoring is not automatically better. An alert that fires on noise trains the on-call engineer to ignore alerts, which is worse than having no alert at all — it means the one that matters gets dismissed along with the rest. Every alert should map to a condition a human can act on, not a metric that happened to cross a threshold. “Error rate above 5% for five minutes” is actionable. “CPU is at 80%” usually isn’t, unless it reliably precedes something that breaks.
The same discipline applies to dashboards. A dashboard with forty panels doesn’t get checked during an incident — it gets ignored in favor of whatever three numbers the on-call engineer already has memorized. A handful of dashboards built around specific failure modes, kept current as the system changes, beat one exhaustive dashboard nobody trusts enough to read under pressure.
The Business Case Nobody Budgets For
Observability work rarely gets prioritized because its payoff is invisible until the moment it isn’t. The return isn’t measured in a sprint demo — it’s measured in mean time to resolution during the incident that would otherwise have taken four engineers three hours of log-grepping to diagnose, and instead took one engineer twenty minutes because the trace showed exactly which service and query were responsible. That gap compounds: faster diagnosis means shorter outages, which means fewer customers who notice, which is the actual metric leadership cares about even when instrumentation is the line item they’re reluctant to fund.
The teams that get this right don’t treat observability as an incident-response tool. They treat it as a permanent part of how the system is built, revisited every time a new service or dependency is added — because the next incident is never in the part of the system that was well instrumented for the last one.
PNK WORKS builds and instruments backend systems with observability designed in from the start, not bolted on after the first outage. Start a project.
Ready to work together?
Start a Project →