Your system failed. You have no idea what happened. That is the real failure.
Every other failure mode has a fix once you find it. FM11 — Observability Blindness is the failure mode that prevents you from finding anything. A cascading failure you cannot see is a cascading failure you cannot stop. A data corruption you cannot trace is a corruption that keeps spreading.
FM11 sits underneath every other FM. It decides how long each one gets to run.
The 3:14 AM page
At 3:14 AM, the on-call engineer gets a page. Latency P99 is above the SLO threshold. She opens the dashboard.
The graph shows P99 spiked from 120ms to 4,200ms starting at 3:11 AM. Error rate is normal. Traffic volume doubled at 3:09 AM because a nightly batch job kicked off.
She opens the trace for a slow request. The trace shows 3,900ms of the 4,200ms went into a single database call. She opens database metrics. The connection pool is exhausted.
Three minutes after the alert, she has the root cause. The batch job saturated the connection pool. She raises the pool limit. Latency recovers. Total time to resolution: seven minutes.
Now imagine the same incident without observability. She would not have known latency spiked. She would not have known when the batch job started. She would not have known which service call was slow.
The seven-minute incident takes hours. Users find it before she does.
That gap is FM11.
The three pillars
Observability is built from three data types. Each answers a question the other two cannot.
Metrics are numeric aggregations over time. They answer: what is happening now, and how does it compare to before? Rate, error percentage, latency distribution.
Logs are structured events with context. They answer: what specific thing happened, to which entity, at what time, with what result?
Traces are correlated request paths across services. They answer: for this specific request, which services were involved, in what order, and how long did each take?
None of the three is sufficient alone. Metrics tell you P99 is 4,200ms. Metrics cannot tell you which service is slow. Logs tell you a query took three seconds. Logs cannot tell you the full call chain. Traces show the call graph for one request. Traces cannot tell you how many other requests are slow.
Diagnosis is the correlation. An alert fires on a metric. You query traces from that window. You filter logs by trace ID. The three pillars form a pipeline from symptom to cause.
What each pillar costs to get wrong
Metrics have three types. Counters go up only. Gauges rise and fall. Histograms hold distributions.
Latency belongs in a histogram, not a mean. Mean latency of 50ms hides a P99 of 2,000ms. The user who hit the P99 request did not experience 50ms. They experienced two seconds. SLOs are written against P99 for this reason.
Logs must be structured. An unstructured line — "User 123 checked out order 456 at 14:30" — requires regex to parse. A structured JSON log with event, user_id, order_id, trace_id, duration_ms becomes a queryable table. SELECT * FROM logs WHERE event='checkout_completed' AND total_amount > 1000 works on the second. It cannot be written against the first.
Traces need context propagation. Each service adds a span. Spans link via shared trace and span IDs. A trace is a tree. The root span is the entry point. The critical path is the longest chain of dependent calls. That chain determines total latency.
Miss any of these disciplines and the pillar becomes decoration.
Where FM11 actually appears
The uninstrumented service. A new service ships without metrics. It handles 20% of checkout traffic. Three months later, P99 checkout latency is elevated. The dashboard shows every monitored service as healthy. The uninstrumented service is the bottleneck. It is invisible.
Engineers spend days diagnosing the wrong services. Someone finally notices the new service has no metrics. The fix is a deployment policy: no service reaches production without a minimum set — request count, error rate, latency histogram.
Broken trace context propagation. A service accepts an incoming request but does not forward the traceparent header. Downstream calls start new traces, unlinked from the original. The root span is visible. The downstream spans are orphans. The complete request path is unreadable. Frameworks that propagate context automatically — the OpenTelemetry SDK is the standard — prevent this without application code changes.
Missing coverage for async paths. A request triggers a Kafka event. A worker consumes it. The HTTP span finishes before the worker runs. Without trace context in the message headers, the worker's processing is invisible in the original trace. Async tracing requires injecting the trace ID into the payload or headers and using it in the consumer.
Each of these is FM11 in a different disguise. In each, some part of the system runs without leaving evidence.
The tradeoff nobody wants to name
The core tension is AT9 — Correctness vs Performance. Collecting 100% of traces gives complete coverage. Every error, every slow request is captured. The cost is storage and network overhead proportional to traffic.
At 10,000 requests per second with ten spans each, that is 100,000 span records per second. Most of them cover fast, successful requests with no diagnostic value.
Head-based sampling decides at the start of the request. Sample 1% of normal traffic. Cost drops by 100×. Rare errors at 0.1% may never appear in the sample.
Tail-based sampling decides at the end. Keep every trace with an error. Keep every trace above the P99 threshold. Keep 1% of the rest. 100% of errors captured. 100% of slow traces captured. The buffer must hold all in-flight trace data simultaneously.
The same tension appears in logs. INFO in steady state. The ability to raise a specific service to DEBUG during an incident. Running DEBUG in production increases log volume 10 to 50×.
It also causes FM3 — Unbounded Log Volume. A service emitting 50GB per day into storage sized for 5GB fills disk in twelve hours. Log volume must itself be a monitored metric.
Metrics have a cardinality trap. High-cardinality tags — user_id, order_id — create one time series per unique value. Ten million users times 100 metric names is one billion time series. Prometheus cannot hold it. High-cardinality analysis belongs in logs and traces, not metrics.
The signal that tells you this applies to your system
A P99 latency alert fires. No service in the dashboard shows elevated latency.
That gap is the signal. Something on the request path is slow. The monitoring cannot see it. Either a service is uninstrumented, or trace context is not propagating, or an async path has no coverage.
You are already in FM11. The alert only told you the symptom. It did not tell you where. Every minute until you find the invisible component is a minute the failure keeps running.
The uncomfortable question
The median time to detect an outage without observability is measured in minutes to hours. With observability, it is seconds to minutes. Mature observability infrastructure costs 10 to 30% of production infrastructure.
That is the deal. You spend a real fraction of your compute and storage budget on watching yourself. In return, incidents that would have surfaced through angry users surface through your own alerts. FM11 shrinks from a permanent condition to a bounded window.
The harder question is the one this article did not answer. When a metric, a log, and a trace disagree about what happened, which one do you trust?
The full framework treatment — compression blocks, three-level exercises, and the complete AT/FM mapping — is in the Reference Book, Chapter 7 (Failure Modes, FM11). Free chapter available at computingseries.com/books/ref.