Mid-senior engineers preparing design rounds

System Design Framework

System design is scored on process, not memorized architectures. This is the five-step framework the MockWise report grades against: clarify, estimate, high-level design, deep dive, trade-offs, with timeboxes and the checklists interviewers listen for.

  • 5 repeatable steps
  • Timeboxes per step
  • Checklist per component

The five steps as one narrative

The framework is not five phases to survive. It is one argument: requirements create the numbers, the numbers choose the architecture, the architecture hides one hard component, and the hard component is where you show depth. Interviewers grade the through-line: can every box, index, and queue be traced back to a number you said out loud in minute three?

  • Clarify (5 min): functional + non-functional + explicit scope cut. Deliverable is a spoken sentence the interviewer can repeat.
  • Estimate (3 min): QPS from daily actives x peak multiplier, storage per year, bandwidth, hot vs cold paths. Deliverable is the number that will justify the next two steps.
  • High-level design (10 min): boxes that each own one piece of complexity, narrated data flow, one justification per box. Deliverable is a topology your estimates explain.
  • Deep dive (10 min): the one load-bearing component to three levels. Schema, keys, partitioning, hot paths, failure and recovery. Deliverable is operational detail, not vocabulary.
  • Wrap and trade-offs (2 min): what your design gives up, where it breaks, what you would revisit first. The sentence that separates passing from strong.

The estimation kit (workable numbers)

  • Powers of two to say confidently: 10^6 seconds ≈ 11.5 days; 1KB x 100k users ≈ 100MB; a modest 100MB/s disk and a 1Gbps pipe; 10^5 is the safe “round number” anchor for anything you invent.
  • DAU to QPS: daily actives / 86,400 x writes-per-user, then a 3-5x peak multiplier and an event-day 10x. Say which one you are using and why.
  • Storage: rows x bytes x retention, plus one backup copy, and name whether the numbers are guesses you are comfortable defending.
  • Latency anchors that make caching arguments land: memory lookup sub-millisecond class, single-digit ms in-datacenter, tens of ms cross-region, hundreds for a third-party call. Round everything and state that you rounded.

Choosing the deep dive

  • Pick the component the constraint actually threatens: fan-out if you have celebrities, the query if you have search, storage if you have write volume. The load-bearing wall, not the interesting one.
  • Signal the switch explicitly (“the hard part is delivery fan-out, so I will go deep there and keep the UI read path at HLD.”) Selection is itself graded, and narration makes it visible.
  • Three levels down, always: what is stored (schema), how it is addressed (keys, indexes), and what happens when it dies (failover, backfill, reconciliation).
  • If the interviewer pushes to another component, drop gracefully with one summary sentence. Attachment to your own deep-dive choice reads as rigidity.

The failure-and-observability checklist

  • Per dependency one line: how you detect (metric, alert), how you degrade (queue, stale read, retry with jitter, circuit open), and how you recover (replay, backfill, reconciliation job).
  • Queue honesty: what happens at 10x backlog. Priority, TTL, poison messages to DLQ, and who gets paged when lag crosses the product budget.
  • The four signals to volunteer unprompted: latency percentiles, error rate, saturation, and data freshness, with one sentence on what each page means at 3am.
  • Consistency language that stays honest: name what a user could observe mid-failure (stale counts, duplicate-tolerant views, eventual feeds). Never say “strongly consistent” to sound senior.

HLD patterns with their cost sentences

Cache-aside vs write-through

Cache-aside: simplest, staleness window accepted, misses carry the DB. The cost sentence is invalidation discipline plus cold-start behavior after deploys.

Queue-based fan-out

Decouples ingest from delivery and absorbs bursts. The cost is ordering (per-key at best), at-least-once duplicates, and the lag budget you now own.

Shard by the hot key

Scales the dominant access path. The cost is cross-shard features: joins, transactions, and global counters become application problems you must design deliberately.

CDN plus short TTL for reads

Turns a DB into an afterthought for hot reads. The cost is the staleness contract; state which user-visible numbers are allowed to be a minute behind.

Practicing the framework

  • Step one only: take three prompts a day and spend five minutes clarifying. The muscle that fails first is stopping before designing.
  • Full 30-minute single-prompt sessions for the argument flow; 60-minute double-prompt for selection and switching under time.
  • Same prompt twice in a week: first pass builds, second pass fixes only the missed points the report named.
  • Keep the five-item round checklist (API, data model, scaling bottleneck, failure, monitoring) as a verbal pre-flight, and audit one skipped item per session.

Example questions and ideal structure

Beginner

Step 1: Clarify before you draw.

Graded on constraint discovery: functional scope, scale, read/write ratio, latency and consistency needs, data lifetime, and an explicit cut of what you will not design.

Why it matters: Panels grade minute two hardest because everything later inherits it. A wrong assumption surfaces in the deep dive as an unexplainable box.

Common mistake: Treating clarification as politeness questions (“is mobile okay?”) instead of numbers and invariants.

Trade-off: Asking too much burns the clock; five sharp questions, then move. The wrap-up confirms scope once more.

Ideal structure: Functional + non-functional + constraints: users, scale, read/write ratio, latency, consistency, and what you will NOT design. Confirm the scope in one sentence before any box.

Tip: Designing before clarifying is the most common rejection reason.

Follow-up to expect: Expect one requirement flipped mid-interview (“it is actually 100x write-heavy.”) The grade is which of your boxes visibly survives it.

Intermediate

Step 2: Estimate, then choose your deep dive.

Estimates exist to justify later choices: QPS drives the queue, size drives the shard, hot/cold split drives caching. Interviewers listen for the callback, not the arithmetic.

Why it matters: The bridge that makes the rest look reasoned; without it every component reads as taste.

Common mistake: Precision theater. Three significant figures with no derivation, or numbers that compute once and echo never.

Trade-off: Time: three minutes, rounded, aloud, with the 10x event-day spike named; skipping it feels fast and costs everything downstream.

Ideal structure: QPS from daily actives (divide, then 3-5x for peaks), storage per year (rows x size x retention), bandwidth. Round numbers, said out loud; then pick the one component where the hard problem lives.

Tip: Estimates justify the architecture; no numbers, no senior score.

Follow-up to expect: “Which number changes your design if I multiply it by 50?” Knowing your sensitive variable is the meta-skill the follow-up checks.

Intermediate

Step 3: High-level design: boxes that own complexity.

Every box must answer “what number forced you,” and the flow must be narratable in one breath: ingest, fan-out, store, serve. Structure is the grade, novelty is not.

Why it matters: It is the section where weak candidates collapse into pattern lists they cannot defend.

Common mistake: Component maxing. Kafka+Redis+K8s+CDN where an app server and Postgres survive the estimate.

Trade-off: State what each box costs (ops, latency, staleness) and which pain it removes; uncosted diagrams are read as tutorials.

Ideal structure: Client → load balancer → stateless service → cache + primary/replica DB, with a queue wherever fan-out exists, and one justification sentence per box.

Tip: Any box you cannot justify is the box they probe.

Follow-up to expect: “What breaks first at double the writes?” Point at the component and the metric that would have warned you.

Advanced

Step 4: Deep dive + trade-offs to the end.

The depth interview inside the interview: schema, keys, partitioning, hot-path behavior, retries, consistency, then the closing sentence naming what you gave up.

Why it matters: The senior signal is choosing the load-bearing component and walking it to the operational floor, not breadth again.

Common mistake: Depth-shopping. Drifting to the fun component (the ML ranking!) while the storage bottleneck nobody justified.

Trade-off: Name the sacrifice per choice: availability for consistency, ordering for throughput, staleness for cost. The sentence is the point.

Ideal structure: One component to three levels: schema, keys, partitioning, hot-spot handling, retries, consistency choice, and finish by naming what your design gives up.

Tip: Say weakly consistent only if you can state the user impact.

Follow-up to expect: “Now the failure walk: that shard dies at peak.” Replica promotion, backlog drain, and what users saw in the gap. The finish is the grade.

Advanced

How would you know this design is failing before users do?

Observability as a design element: the two or three metrics per component that map to user pain, their alert budgets, and the synthetic probe that catches silent failures.

Why it matters: The question that separates people who operate systems from people who draw them. Asked more and more as loops modernize.

Common mistake: Dashboard recitation (“Prometheus and Grafana”) with no metric named against a threshold or a user symptom.

Trade-off: Signal volume vs fatigue. The design ships four pages, not forty, and the answer defends the cut.

Ideal structure: Per component: the golden signals plus freshness/lag for async paths, the 3am page rule, and the one synthetic transaction that catches what graphs miss.

Tip: End with what would be un-monitored and how you accepted that risk. Complete honesty scores.

Follow-up to expect: “A silent corruption: writes succeed, values wrong. Detection?” Sampling checks and reconciliation jobs are the expected answer.

How to prepare

  • Timebox it: 5 min clarify, 3 estimate, 10 high-level, 10 deep dive, 2 wrap. The interviewer is timing you too.
  • Run a five-item checklist per round: API, data model, scaling bottleneck, failure handling, monitoring. Miss one and the report marks it.
  • Do the same prompt twice: once to build the frame, then 30-minute drills on your two weakest boxes.
  • Say every estimate out loud and callback to it. “this is why the queue exists” is the single sentence reviewers listen for most.
  • Keep three deep-dive archetypes rehearsed (storage schema, fan-out, cache layer); novel prompts map onto one of them in the first three minutes.

FAQs

Which design questions actually appear?

Real loops mix classics (URL shortener, rate limiter, notification fan-out) with open-ended scale prompts. The framework transfers; memorized solutions do not.

How does MockWise grade this?

The report scores Requirements, Estimation, HLD, and deep-dive quality, and lists missed points (like no backpressure discussion) as drills to redo.

Do I need to practice drawing if the session is verbal?

The verbal version is the harder, more transferable skill. “top-left to bottom-right, four boxes, here is each one’s job” is exactly what phone and on-site loops require when you cannot share a canvas.

How many distinct prompts should I have done?

Eight to ten across the families (fan-out, storage, search, real-time, rate limiting) with two passes each beats thirty skims. The framework only proves itself through repetition.

Is low-level design part of this?

Design rounds stay high-level; service-level LLD (class and function structure, error handling) is a different round. For LLD-heavy loops the technical sessions’ API questions are the closer drill.

Ready to practice?

Practice your answers in a mock interview, then review the transcript-based feedback.

Related practice

Last updated: 2026-09-03 · Questions? Contact support · Security