Reliability & operations

Observability

Observability is designing a system so its current state can be understood from outputs. Correlating metrics, logs, and traces to answer “what is happening and to whom,” including questions you did not predict.

In technical terms

The three pillars: metrics (aggregates, thresholds, SLOs and error budgets), structured logs (per-event context, redaction for PII), distributed traces (spans per request, sampling to afford them). Golden signals (latency, traffic, errors, saturation) map to user pain; alerting works best on symptoms (user-facing SLO burn) rather than each cause.

Why it appears in interviews

Design rounds end with “how would you know it is failing?” because the answer reveals whether you have operated what you draw. Unmonitored architectures cap the score no matter how clean the boxes look.

The common misconception

Dashboards are not observability. Walls of green graphs you never query answer known questions; the property under test is novel ones: “which users on checkout hit the one broken region, traced to which dependency span?”

Trade-offs & when it hurts

Full-fidelity tracing costs storage and CPU (tail-based sampling buys honesty on the weird requests); verbose logs leak PII without redaction discipline; cause-based alerting generates fatigue because the page must mean something. Ship the three sentences: what I would watch, what I would page on, and what I accept unseen.

How to show it in an interview

Close every design with the page list: “p99 latency, error rate, queue age, replica lag. Alerts on SLO burn over five minutes, not CPU; the trace query I would write first is the one that finds a slow request end to end.” Four signals, one rule, one query: monitoring stops being a box on the diagram and becomes operation.

Questions this concept earns

  • A deploy is going wrong for 3% of users in one region: which three views answer that?
  • Silent corruption: writes succeed, values wrong. What exists in your design to catch it?
  • Your paging threshold: what gets a person up at 3am, what is a dashboard, and what is invisible? Defend the line.

Use the concept in a real session

Answer follow-up questions about observability and related systems, and get a scored report in minutes.

Related

Last reviewed: 2026-09-03 by MockWise Engineering · Corrections welcome via contact.