Data scientists, analysts, ML engineers

Data Scientist Mock Interviews

DS loops test SQL + modeling judgment + product sense. MockWise probes your portfolio with metrics-first follow-ups and scores “metric framing” alongside accuracy.

  • Offline vs online framework
  • Portfolio with impact numbers
  • SQL with ties + performance

What DS loops actually grade

  • Judgment over tooling: which metric you chose, why the baseline came first, and where your own numbers might be lying. Frameworks are assumed, skepticism is graded.
  • Metric framing as the default grammar: every claim arrives with measure, direction, and cost; “improved accuracy” without a baseline or business number is the round’s signature failure.
  • SQL as the gate round: correct grain, tie handling, window functions, and reasoning about plans. Analysts and DS alike get cut here more than in modeling.
  • Communication to non-technical stakeholders: being able to explain the model, its limits, and the decision it enables to a PM inside sixty seconds.
  • Product sense on your own work: did the project matter to anyone, how did you know, and what would you build differently knowing what you know now.

The project walkthrough, station by station

Problem framing

The business question before the ML question: whose behavior changes, by how much, worth what. Interviewers listen for whether you can state success before describing the model. Most candidates never do.

Data story

Grain, size, and mess: what the rows are, what you excluded and why, leakage you hunted (future info, target-derived features), and the join fan-outs that create duplicate-label traps. Provenance questions (“how did you know the labels were right?”) probe this station hardest.

Modeling judgment

Why the baseline first, why this model over the obvious alternative, what it costs. Complexity, latency, interpretability you gave up. The grade is the sentence naming what you sacrificed, not the model name.

Impact and honesty

Offline metric, online test (or why none and what you used instead), the business number it moved, one thing you would do differently, and the moment the project lied to you. The last two are the credibility stations. Skip them and the walkthrough reads as marketing.

The offline-vs-online ladder

  • Offline metrics measure ranking quality on logged data. NDCG/Precision@k with time-cut splits; they are cheap, fast, and quietly full of lies: leakage, feedback loops, and popularity bias all inflate offline numbers.
  • The honesty check is naming one inflation mechanism per metric you quote. Interviewers grade that sentence more than the metric name itself.
  • Online means an A/B with a primary success metric, pre-registered guardrails (latency, complaint rate, revenue per session), and a ramp plan with a rollback trigger.
  • Between them sits calibration and stability: offline lift that does not survive into online lift is normal and expected. Having a hypothesis for why shows senior judgment.
  • Post-launch: drift monitoring (feature distributions, label lag, population shift) with a re-evaluation cadence; unmonitored models are how re-interview questions get harder.

SQL the way DS rounds ask it

  • Cohorts and retention: define the cohort, window by signup week, and join behavior back to it. The grade lives in the grain statement before any code.
  • Funnels and windows: COUNT(DISTINCT user) traps, ORDER BY with PARTITION BY for rank/running math, and naming the difference between ROW_NUMBER and DENSE_RANK on ties.
  • Deduplication: same event from multiple sources. Rank by freshness and filter to one row per key; interviewers always check NULL and tie semantics.
  • Performance reasoning: which index helps this query, what EXPLAIN would show, why window functions beat correlated subqueries at 10M rows. You do not execute here, so say the plan.

How MockWise scores DS sessions

  • Metric framing is detected as a habit: answers that name measure, baseline, and business number score separately from answers that name models.
  • Portfolio probes arrive metrics-first (“what did it move, how do you know?”), and vague impact claims become named missed points you can re-drill.
  • SQL questions grade reasoning: grain, tie handling, index and plan commentary. No execution required; follow-ups change constraints mid-question and watch you adapt.
  • The sixty-second model explanation is graded as its own dimension: communication scores whether a PM could repeat your answer back.

What sinks strong modelers

  • Baseline skip: going straight to the fancy model and forfeiting the “what did it beat?” question that grades the whole story.
  • Metric worship: quoting AUC or accuracy with no cost asymmetry. The fraud-threshold variant exists precisely to catch this.
  • Unverifiable impact: numbers with no measurement story behind them; provenance questions always come.
  • Silent assumptions: no leakage check named, no grain declared. The follow-up finds them; better when you do.
  • Vocabulary without experience: naming XGBoost or embeddings without one sentence of what broke last time you used them.

Example questions and ideal structure

Advanced

How do you evaluate a recommendation model offline vs online?

Offline metrics (NDCG, Precision@k on time-split logs) rank candidates cheaply but inherit feedback-loop bias: you mostly evaluate on what the old system already showed users.

Why it matters: The eval question is where interviewers separate people who shipped models from people who trained notebooks.

Common mistake: A big offline lift proving real-world lift. Position bias and exposure skew routinely eat 30-50% of the gain.

Trade-off: Online wins trust but costs latency of iteration: name the guardrails and ramp speed you accepted as the price.

Ideal structure: Offline (NDCG, Precision@k, leakage checks) vs online (A/B, guardrails, ramp), and why offline can lie, with your own project as example.

Tip: Name the leakage check you ran.

Follow-up to expect: “Offline +12%, online +2%: list three plausible reasons”: exposure bias, novelty decay, and metric mismatch are the expected trio.

Intermediate

Walk me through a recent data project end-to-end.

Not a tour. An audit trail: business question, data grain and mess, baseline, modeling choice with its cost, eval results, shipping decision, measured impact, and one honest regret.

Why it matters: It is the single most predictive DS interview question; everything the loop cares about (judgment, communication, provenance) shows up here.

Common mistake: Optimizing the story’s ambition over its traceability. Panels grade the second follow-up on your data story, not your headline.

Trade-off: Include the trade-off you actually made (complexity vs interpretability, speed vs rigor). Its absence reads as inexperience.

Ideal structure: Problem → data → baseline → modeling → eval → ship → impact (e.g., +3% CTR, -$12k/month), with a trade-off you made.

Tip: Quantify impact, not just accuracy.

Follow-up to expect: “What would you change today?” A specific upgrade (new label source, calibration layer, simpler model with features) proves the learning is real.

Intermediate

SQL: second highest salary per department, handling ties, and make it fast on 10M rows.

The per-department turn makes PARTITION BY mandatory, and “second highest value” versus “second row” is DENSE_RANK versus ROW_NUMBER. State which you are answering first.

Why it matters: Grain + semantics + performance in one question is exactly how DS SQL screens are written.

Common mistake: The two-MAX subquery still technically works for values but collapses into per-row scans without an index and confuses NULL departments.

Trade-off: Window scan (one pass, needs (dept, salary) index) vs self-join (reads twice). Name which plan EXPLAIN shows and why.

Ideal structure: DENSE_RANK vs ROW_NUMBER, window vs self-join, index on (dept, salary), and the tie semantics you chose.

Tip: State tie handling before writing SQL.

Follow-up to expect: “Only departments with 5+ employees, top 3 salaries.” HAVING versus WHERE placement on the rank filter is the expected follow-up.

Advanced

How would you set a threshold for a fraud model with 99% true negatives?

Accuracy is meaningless at that imbalance. The decision moves to the cost matrix: missed fraud loses money, false declines lose customers, and the threshold sits where the marginal costs cross.

Why it matters: It tests whether you optimize business outcomes or leaderboard metrics; every production DS answer needs the cost sentence.

Common mistake: Tuning for best F1 by default. F1 bakes in equal costs you have not earned or checked.

Trade-off: Name whose review queue absorbs the false positives and at what labor cost. The threshold is an operations decision wearing a model hat.

Ideal structure: Precision-recall vs cost matrix, threshold via expected cost, and monitoring for drift.

Tip: Bring a cost number.

Follow-up to expect: “Approval rates drop and fraud barely moves: adjust or investigate?” The drift-before-tuning answer is the grade.

Intermediate

We want to A/B test a checkout redesign. Design the experiment.

The structure is the answer: population and unit of randomization, primary metric with a pre-registered threshold, guardrails, minimum detectable effect driving the sample-size math, duration in full weekly cycles, and the stopping rule.

Why it matters: Experiment design appears in DS, analytics, and product-adjacent rounds alike. It grades causal hygiene under practical constraints.

Common mistake: Peeking daily and stopping at significance. Name sequential testing or a fixed horizon instead.

Trade-off: Shorter tests need larger effects; a 2% conversion lift demands days you may not have. State which constraint you traded.

Ideal structure: Hypothesis → metric + MDE + power → unit and splits → duration in cycles → guardrails → pre-registered decision rule.

Tip: Say “minimum detectable effect” once with a real number and the interviewer hears experience.

Follow-up to expect: “The test moves conversion but not revenue: what now?” Metric-hierarchy reasoning: which was the primary and why the mismatch is informative, not broken.

Advanced

You cannot run an experiment on this feature. Did it work?

Quasi-causal reasoning: pre-post with seasonality controls, difference-in-differences across an unaffected segment or region, synthetic control with honest assumptions, and the sentence on what still confounds it.

Why it matters: The most common real job problem (no clean A/B available) and the strongest seniority signal in analytics interviews.

Common mistake: Presenting a trendline as impact; panels specifically probe what would have happened without the feature.

Trade-off: Every quasi-experiment trades internal validity for feasibility. Naming which assumption your result hangs on is the pass condition.

Ideal structure: State the counterfactual you cannot get, the best available proxy comparison group or time control, its threats, and the metric plus guardrails you still measure.

Tip: Volunteer the confounder interviewers would use to check you. It is the whole grade.

Follow-up to expect: “What evidence would make you downgrade the claim to inconclusive?” Pre-committing to a falsifier is the senior move.

How to prepare

  • Every answer: metric + baseline + trade-off. "We improved accuracy" is not enough. Name the metric and the cost.
  • Practice explaining one model to a PM in 60 seconds. MockWise scores communication separately.
  • Review “missed points” for SQL and redo until “handling ties” and “index” are mentioned without prompting.
  • Build the portfolio audit before the mock: for each project, one business number with provenance and one honest regret. The walkthrough is scored on both.
  • State data grain and tie semantics out loud before any SQL logic; it is the most commonly missed point in DS rounds.

FAQs

Do I need to write SQL here?

Explain the query verbally. MockWise scores correctness, tie handling, and performance reasoning, not execution.

How is portfolio scored?

We extract impact numbers; reports penalize vague “we built a model” without a metric.

Which interview type should I run?

Technical sessions carry SQL, modeling, and case probes with data follow-ups; pair one with a behavioral session before loops. The project walkthrough is graded as both.

Is this for analysts too or just DS?

The SQL-cohort-communication core is analyst-round material; model-eval and experiment-design stations deepen toward DS. The AI calibrates depth to the resume you upload.

What if my best project had no business impact?

Grade it honestly and show the reasoning loop: what you measured, what you would instrument now, and what the result taught. Panels forgive small numbers far more readily than invented ones.

Ready to practice?

Practice your answers in a mock interview, then review the transcript-based feedback.

Related practice

Last updated: 2026-09-03 · Questions? Contact support · Security