Performance

p99 / Percentile Latency

Percentiles describe a response-time distribution: p99 means 99% of requests completed at least that fast. The unlucky 1% are the users filing the bug report.

In technical terms

Tails come from the compounding of rare events: GC pauses, lock contention, queueing when a resource saturates, cold caches, retry storms, and fan-out to the slowest of N dependencies. Means hide tails (2% timeouts with a great average); percentile-of-percentiles noise makes naive p99s of p99s wrong. Aggregate histograms (HDR, t-digest), never average percentages.

requests   p50    p95    p99    max
2.1M       41ms   220ms  1.4s   9s   <- 1% of users see 34x the typical

Why it appears in interviews

SLA questions, capacity questions, and the classic debugging prompt all speak percentiles; misusing “average” in either direction reads as never having owned a latency budget.

The common misconception

That the mean plus a count is enough, and that p99 across services is the sum of p99s: tails are not independent like that, and the distribution shape (bimodal? periodic?) carries the diagnosis.

Trade-offs & when it hurts

Squeezing the tail is costly. Hedged requests pay bandwidth, warm capacity pays idle spend, timeout budgets convert tail into errors. Define the user-observable budget first; optimizing a percentile nobody perceives is dashboard theater.

How to show it in an interview

Put the percentile in a budget: “The page gets 2.5 seconds; that leaves 800ms p99 for this API, so I hedge the read after 600ms across two replicas, accepting 2x fan-out cost to protect the tail the user actually feels.” Number, budget, source of the tail, and what it costs: the complete sentence.

Questions this concept earns

  • p50 is flat, p99 tripled after deploy. The first three things you check, in order.
  • Why is averaging per-service p99s mathematically wrong, and what aggregates instead?
  • Hedged requests cut p99 but raise load 30%: where do you stop hedging and why?

Use the concept in a real session

Answer follow-up questions about p99 / percentile latency and related systems, and get a scored report in minutes.

Related

Last reviewed: 2026-09-03 by MockWise Engineering · Corrections welcome via contact.