Performance
p99 / Percentile Latency
Percentiles describe a response-time distribution: p99 means 99% of requests completed at least that fast. The unlucky 1% are the users filing the bug report.
In technical terms
Tails come from the compounding of rare events: GC pauses, lock contention, queueing when a resource saturates, cold caches, retry storms, and fan-out to the slowest of N dependencies. Means hide tails (2% timeouts with a great average); percentile-of-percentiles noise makes naive p99s of p99s wrong. Aggregate histograms (HDR, t-digest), never average percentages.
requests p50 p95 p99 max
2.1M 41ms 220ms 1.4s 9s <- 1% of users see 34x the typicalWhy it appears in interviews
SLA questions, capacity questions, and the classic debugging prompt all speak percentiles; misusing “average” in either direction reads as never having owned a latency budget.
The common misconception
That the mean plus a count is enough, and that p99 across services is the sum of p99s: tails are not independent like that, and the distribution shape (bimodal? periodic?) carries the diagnosis.
Trade-offs & when it hurts
Squeezing the tail is costly. Hedged requests pay bandwidth, warm capacity pays idle spend, timeout budgets convert tail into errors. Define the user-observable budget first; optimizing a percentile nobody perceives is dashboard theater.
How to show it in an interview
Put the percentile in a budget: “The page gets 2.5 seconds; that leaves 800ms p99 for this API, so I hedge the read after 600ms across two replicas, accepting 2x fan-out cost to protect the tail the user actually feels.” Number, budget, source of the tail, and what it costs: the complete sentence.
Questions this concept earns
- p50 is flat, p99 tripled after deploy. The first three things you check, in order.
- Why is averaging per-service p99s mathematically wrong, and what aggregates instead?
- Hedged requests cut p99 but raise load 30%: where do you stop hedging and why?
Use the concept in a real session
Answer follow-up questions about p99 / percentile latency and related systems, and get a scored report in minutes.
Related
Last reviewed: 2026-09-03 by MockWise Engineering · Corrections welcome via contact.