Skip to content
Question Vault?
Free to readNo accountNo email wallNo invented statisticsNo partial listsCopy or print any set and take it with you

Questions to Ask in a Data Science Interview

For data scientists and data analysts who have reached the point in an interview where the floor is handed over. The questions follow the order in which you would size up the job: what the team works on, the state of the data and tools, how a model reaches production, who asks for the work and how it is prioritized, the team and where the role leads, and a few to close on. Each has a note on what a solid answer sounds like, what a worrying one sounds like, or what to do with it.

54 questions

Want questions from the whole vault instead? Try the random question generator.

The questions

Each question, and why to ask it

The work

What kinds of problems does the data science team work on, and which one would I start with?

Why ask it

A strong answer names a problem with a decision attached, such as which customers get a retention offer or how much stock a warehouse orders. 'We want to get more out of our data' suggests there are people and no agreed problem yet, and your first months could go to finding one. Ask what the last person hired began on.

What does a typical week look like for a data scientist on this team?

Why ask it

Put this to someone who does the job if you can, and ask about last week in particular. Sort what you hear into time in SQL and notebooks, time in meetings and time answering other people's requests. A week that is mostly pulling numbers on demand is a reporting job, whatever the title says.

How is the role split between analysis, modeling and data engineering?

Why ask it

One title covers dashboards and experiment readouts at one company and production machine learning at the next, so ask for rough proportions. Then ask whether the split is expected to shift after the first year. None of the mixes is wrong, but only one of them is the work you want to get better at.

What is the biggest challenge the data science team is dealing with right now?

Why ask it

The answer usually lands in one of three places: the data is missing or untrusted, finished work does not reach production, or nobody acts on the results. Whichever it is, that is the problem you would be hired into, so spend your next question on it. Be a little wary of 'no real challenges', and put the same question to a peer on the panel if you meet one.

What would you want a new data scientist to have delivered by the six-month mark?

Why ask it

Good answers are things you could point to: a forecast the planning team relies on, an experiment read out, a model serving live traffic. Compare the size of it with what you have heard about the data. A production model in six months on tables nobody trusts is a goal that was set without checking the ground.

What did the team deliver last quarter that changed a decision or a product?

Why ask it

You are listening for the second half: a price that moved, a feature that shipped, a campaign that was dropped. A report that was well received does not count until you have heard what anyone did differently after reading it.

Why is the team hiring a data scientist now?

Why ask it

Growth, a departure and a brand-new function each hand you a different job. If someone left, ask what they were working on and whether it is waiting for you. If the function is new, ask who pushed for it and what they told leadership it would deliver.

How much of the work is building something new, and how much is keeping existing models and pipelines running?

Why ask it

Mature teams spend a real share of their time on upkeep, and an honest interviewer will say so. Be more wary of 'all new work' from a team that also claims many live models, since someone is maintaining them. Ask whether that someone would be you.

Is the team working with large language models, or is most of the work classical statistics and machine learning?

Why ask it

A posting can mention generative AI without the team having shipped anything that uses it. Ask which project relies on it today and who that project serves. If the day-to-day is regression, forecasting and SQL, that can be a perfectly good job, and you should choose it knowing that.

Data and tools

What state is the data in: could I query clean tables on day one, or does every project begin with cleanup?

Why ask it

Nobody's data is tidy, so a claim that it is all in great shape deserves a follow-up about the worst table. The reassuring answer admits the mess and names who is fixing it. A useful number to come away with is how long the last project spent on preparation before any analysis began.

Which tools does the team use day to day: Python or R, which warehouse, which cloud?

Why ask it

The list matters less than how settled it is. Ask what the team is moving away from, because a half-finished migration means working in two systems for a while. If the stack differs from yours, find out how long they expect a new person to take to get comfortable, and whether that time is allowed for.

Is there a data engineering team, and where does their work stop and a data scientist's begin?

Why ask it

Where there are data engineers, find out how long a request for a new table or a fixed pipeline usually waits. Where there are none, you are the data engineer, and the modeling share of the job shrinks to match. Either can suit you, as long as you hear it before you accept.

Is there one agreed definition of the main metrics, such as an active user or a churned customer?

Why ask it

Without one, meetings go to arguing over whose number is right instead of what to do about it. A good sign is a written set of definitions, or a shared metrics layer, that a named person owns. The practical test is two dashboards that disagree: find out who settles it.

When a pipeline breaks or a source table changes, how does the team find out?

Why ask it

Automated tests and alerts mean someone has invested in trust. The worrying version is a stakeholder's email, because by then a wrong number has already been shown to someone. Have them describe the most recent incident, including how long the bad data sat there before anyone noticed.

For the first project, is the outcome I would be predicting already recorded in the data, and how far back does it go?

Why ask it

Plenty of projects stall because the label does not exist: churn was never defined, fraud was never confirmed, or history starts last spring. A manager who has checked can answer in a sentence. 'We assume so' means the first phase of the project is finding out whether it is possible at all.

If the data needed to answer a question is not being collected, how long does it take to start collecting it?

Why ask it

This shows whether data scientists have any pull with the engineers who control logging. A matter of days or one sprint means tracking is treated as part of the product. A quarter and an escalation means answering questions with whatever happens to exist.

Is there documentation for the tables and their quirks, or does that knowledge live in a few people's heads?

Why ask it

A data catalog that is kept current shortens your first months considerably. If the answer is a person's name, ask how much of their time you could have, and what happens when they are on vacation. Offering to write things down as you learn them is usually welcome.

Does anyone review analysis and model code before the results are relied on?

Why ask it

Review catches the join that silently doubled the rows and the leak that made a model look brilliant. Ask whether notebooks go through version control and who last reviewed the interviewer's own work. Unreviewed work still gets checked eventually, by the stakeholder who spots the wrong number.

What compute would I have for training models: a laptop, a shared cluster or cloud GPUs?

Why ask it

Skip this if the role is analytics with no model training. Otherwise the answer sets the ceiling on what you can try, so find out who approves the spend as well. A telling follow-up is whether anyone has had an experiment refused on cost.

How does the team handle personal or sensitive data, and who decides what may go into a model?

Why ask it

The rules depend on the country, the industry and the company's own policies, so ask how it works there instead of assuming. What you want is a named owner and a known process, such as a privacy review before a new data source is brought in. A shrug means those decisions would fall to you.

Production

How does a model get from a notebook into production, and who does that work?

Why ask it

There are three common answers: the data scientist ships it, a machine learning engineer rebuilds it, or it has never quite happened. Ask them to walk through the most recent one step by step. Hesitation over the steps usually means each launch is improvised.

How many models built by this team are running in production today?

Why ask it

A plain number, with what each one does, is the answer of a team that ships. Zero is not disqualifying for a young team, but ask what has stood in the way and whether this hire is meant to change that. Count scheduled forecasts and scoring jobs too, since those are production in every way that matters.

How long did the last model take to go from first idea to live use?

Why ask it

Weeks suggests working infrastructure and a team that scopes small. Many months is common where data access, review and engineering queues all sit between you and launch. Ask where the time went, because the slowest step is where you would be waiting too.

How do you decide a model is good enough to launch: an offline metric, an A/B test, or someone's sign-off?

Why ask it

Best is a bar agreed before the work starts and a live test against what exists today. Be wary when the only gate is accuracy on a held-out set, since that says little about whether the business result moves. Find out who can say no, and whether they ever have.

Is there an experimentation platform, and who decides when an A/B test is called?

Why ask it

Where tooling exists, ask who sets the sample size and the stopping rule. If product managers end tests when the number looks good, you will spend energy defending basic statistics. With no platform at all, the thing to learn is how the last launch was measured.

When a simple baseline does almost as well as a complex model, which one ships?

Why ask it

Teams that have maintained models for a while tend to pick the simple one, and can tell you about the time they did not. An answer that favors the sophisticated method every time hints that work is judged on how impressive it looks. Notice whether they have a baseline habit at all.

Once a model is live, who watches it, and what happens when its performance drifts?

Why ask it

Look for monitoring on inputs and outcomes, an alert threshold, and a named owner. If the answer is that someone would notice eventually, models are being launched and left. The story of how the last degraded model was caught, and by whom, tells you more than a description of the dashboard.

How often are models retrained, and is that scheduled or done by hand?

Why ask it

A scheduled, tested retraining job means the team treats models as software. Retraining by hand when somebody remembers is fragile, and it is also a project you could offer to take on. The follow-up is what happens when a retrained model turns out worse than the one it replaced.

Is there an on-call rotation for models or pipelines, and would I be on it?

Why ask it

Many data science roles carry no pager, but teams that own live scoring or nightly pipelines often do. Get the size of the rotation, how often your turn comes and what the last out-of-hours alert was about. How on-call time is compensated differs by employer, so ask how it works there.

Can you tell me about a model that failed after launch or never launched, and what the team changed afterward?

Why ask it

Most teams that ship have one. A specific story with a cause, such as training data that did not match live traffic, and a change in practice shows a team that learns. No story at all means either nothing ships or nobody looks back.

Does anyone have to explain a model's decision to a customer, an auditor or a regulator?

Why ask it

In lending, insurance, hiring, health and similar areas the answer can shape which methods you are able to choose. Requirements differ by country and industry, so ask what applies to this team and who checks it. Where explanations are required, it helps to know which kinds of model have been approved so far.

Stakeholders

Who asks the team for work, and how do those requests arrive?

Why ask it

A ticket queue, a planning meeting and a direct message from a vice president are three different working lives. Ask which partner takes the most time and what that relationship is like. Requests that arrive from everywhere with no intake mean your priorities get set by whoever messaged last.

Who decides what the team works on next, and can a data scientist propose a project?

Why ask it

A team that only takes orders works like a service desk, and its people rarely get to choose interesting problems. A team that only follows its own curiosity can drift away from anything the business needs. The example to ask for is the last project that began as a data scientist's idea.

How much of a data scientist's week goes to one-off requests for numbers?

Why ask it

Some of it comes with every data job, and it is a quick way to learn the business. When it takes most of the week, longer projects stop moving, so ask for an honest estimate. A team that rotates the duty, or has analysts who handle it, has thought about protecting project time.

Is the team brought in before a decision is made, or asked afterward to measure it?

Why ask it

Being in the room early means the analysis can change what gets built. Being called in afterward often means producing a number to support a choice already made. Ask for a recent decision and at which meeting someone from data first heard about it.

Can you tell me about a time an analysis contradicted what a senior person wanted to do?

Why ask it

The ending is what counts. If the plan changed, or was at least delayed for a test, evidence carries weight there. If the analysis was shelved, or the analyst was sent back to look again until the answer improved, expect the same treatment for your own findings.

How do the people who rely on the team's work react to an inconclusive result?

Why ask it

Honest analysis often ends in 'we cannot tell yet'. Stakeholders who accept that, and fund a longer test, make careful work possible. Where a flat result counts as the data scientist's failure, the pressure runs toward finding something, and that is where bad statistics come from.

Who presents results to the people who act on them: the data scientist who did the work, or a manager?

Why ask it

Presenting your own work is how you get known, and how you learn which questions the business cares about. If findings travel upward through someone else, ask whether you would at least be in the room. Junior candidates can ask how soon that changes.

Who builds and maintains the dashboards: this team, the analysts, or a separate BI group?

Why ask it

For an analyst seat this may be most of the job, so ask how many dashboards there are and how many still get opened. For a modeling seat, find out whether dashboard requests would reach you anyway. Where partners cannot pull their own figures, building that self-service layer may turn out to be the real first project.

Which company goal does this team's work feed most directly?

Why ask it

A team whose work is tied to revenue, retention or cost usually finds it easier to argue for data, engineers and headcount. If the interviewer struggles to connect the work to a goal, ask how leadership hears about what the team has done. A team nobody upstairs can describe is hard to defend at budget time.

Team

Is data science a central team here, or are data scientists embedded with product and business teams?

Why ask it

Central teams give you peers, review and a manager who understands the craft. Embedded seats put you close to decisions but can leave you the only data person in the room. In a hybrid, ask who sets your priorities and who writes your review, since those are sometimes two people.

How big is the data team, and how is it split between data scientists, analysts and engineers?

Why ask it

The ratio tells you what you will end up doing yourself. Several scientists and no engineers means building your own pipelines. One scientist among many engineers means being the statistics authority from the first week, ready or not.

Does the person I would report to have a data science background?

Why ask it

A manager from the field can tell a careful analysis from a lucky one and argue for the time careful work takes. A manager from product or engineering may be an excellent advocate and still be unable to check your methods. In that case ask who can.

Would I be the only data scientist on my projects?

Why ask it

Working solo builds range quickly and suits people with some years behind them. Early in a career it means nobody catches your errors or shows you a better approach. If the answer is yes, ask whether there is a regular review with other data scientists, even ones on different teams.

How long does it take a new hire to get access to the data and run a first query?

Why ask it

Access approvals, security training and warehouse permissions can each hold a new person up, and how long that takes differs a great deal between employers. An interviewer who answers in days has watched it happen recently. If it is weeks, find out whether the starter project is chosen so that it does not wait on the slowest approval.

How is a data scientist's performance judged here: models shipped, business results, or feedback from stakeholders?

Why ask it

Each measure pulls behavior a different way. Counting launches rewards shipping whether or not anything improved, and judging by stakeholder praise rewards saying yes. Ask how a project is rated when the careful conclusion was that the idea did not work, since much sound analysis ends there.

What are the levels above this one, and can a data scientist stay technical while moving up?

Why ask it

Ask for a person, not a ladder: who was last promoted and what they had done. Some companies have a staff or principal track for data science and some expect senior people to manage. If the track exists on paper and nobody is on it, treat it as untested.

Does the team set aside time to read papers, try new methods or present work to each other?

Why ask it

A reading group, a regular demo session or a conference budget that people actually spend all count. 'When there is time' usually means there is none. Ask what the interviewer last learned on work hours and how it found its way into a project.

What separates a good data scientist from a great one on this team?

Why ask it

Listen for whether the answer is technical. Many managers will say the great ones frame the question well and explain results plainly, which tells you communication is rewarded. An answer that is all methods and publications points to a culture that leans toward research.

Closing

What do the remaining rounds cover: a take-home, live coding, a case study?

Why ask it

Ask the recruiter or whoever is coordinating, early enough to prepare. Find out the format, the time allowed, the language and whether the take-home is reviewed with you afterward. Processes differ by company, so do not assume the next round resembles the last place you interviewed.

Which part of my experience would you want to see more evidence of?

Why ask it

It gives the interviewer room to voice a doubt while you can still answer it: no production work, a thin statistics background, an unfamiliar industry. Reply with one concrete project and leave it there. If they name nothing, ask what the later interviewers will be probing.

What do you wish you had known about the data or the team before you joined?

Why ask it

People tend to answer this more candidly than a direct question about problems. The reply often names the thing nobody mentions in interviews, such as how long access takes or which system everyone avoids. Best asked of a peer, without the manager present if the schedule allows.

What have you worked on here that you would put at the top of your own resume?

Why ask it

An interviewer who lights up and describes real work with a result is good evidence that such projects exist there. A long pause, or an example from several years ago, is worth noticing. It also ends the conversation on something they enjoy talking about.

When do you expect to make a decision, and is there anything else you need from me before then?

Why ask it

Keep it for last and write down the date. The second half sometimes brings out a request for a code sample, a portfolio link or references, which you can send the same day. If another offer has a deadline, this is the moment to mention it.

How to use your questions in a data science interview

Practical guidance for the conversation itself

Before the interview

Work out which data science job this is

The title covers at least three jobs. A posting heavy on SQL, dashboards and A/B tests describes product analytics. One that lists Docker, APIs and latency describes machine learning engineering. One built around forecasting, pricing or risk describes modeling for a single business function. Decide which you think it is, then open with the questions under The work to confirm it before anything else. For an analytics seat, most of Production will not apply, so put that time into Data and tools and Stakeholders.

Send each group to the person who knows

A recruiter can answer Closing and say what level the role is. The hiring manager knows The work, Stakeholders and Team. A data scientist on the panel is the one to ask about Data and tools and Production, because they live with both. If a product manager or business partner interviews you, turn the Stakeholders questions around: ask what they need from the data team and whether they get it.

Pick the three that could change your mind

Your turn is often five minutes. Decide beforehand which answers would make you turn the job down, such as no route to production if you want to ship models, or no peers if you want to learn from other data scientists. Ask those first and keep the others as spares.

Use what the process has already shown you

A take-home exercise or case study is a sample of how the team thinks. Ask how close its dataset was to the real thing and what the team would have done next with your result. It shows you paid attention, and the answer is a preview of the actual work.

In the room

Ask about the last one

'How do models get deployed?' invites a description of the intended process. 'How did the last one get deployed?' gets what happened. The same swap works for the last experiment, the last broken pipeline and the last project a stakeholder turned down.

Ask for counts the interviewer would know offhand

How many models are live, how many people are on the team, how long the last launch took. These are quick to answer and hard to dress up. Ask them out of interest, not as an audit, and accept an estimate.

Put one question to two people

Ask the manager and a peer how much of the week goes to one-off requests, or how clean the data is. Managers describe the team they are building and peers describe the one that exists. The distance between the two answers is useful in itself.

Keep the data questions friendly

Every company's data has problems, and the people across the table know theirs well. 'What should I expect to run into?' gets a fuller answer than anything that sounds like an inspection.

Reading the answers

Messy data is normal, unowned data is the warning

Do not mark a team down for admitting the tables are a mess. Mark it down when nobody is responsible for fixing them, no time is set aside for it, and the plan is that the new hire will cope.

Listen for decisions, not deliverables

Dashboards built and models trained are outputs. What tells you the team matters is a decision that went differently because of them. If a full round of interviews produces no such example, ask for one directly before you accept anything.

Weigh a first-hire role on its own terms

Being a company's first data scientist can be a fine job or a lonely one. It tends to go better when there is a sponsor with authority, some engineering help with the data, and one concrete problem to begin on. Ask about each of the three, and be cautious if all are missing.

Judge it against the job you want

A role that is mostly analytics is a disappointment to someone set on machine learning and a good fit for someone who likes being near product decisions. Before the interview, write down the mix you are after, and compare the answers with that instead of with a general idea of a good team.

Mistakes to avoid

Asking only about the stack

Tools are the easiest thing to ask about and the easiest thing to learn on the job. Whether the work reaches production, and whether anyone acts on it, will shape your time there far more than the choice of warehouse.

Turning your questions into a quiz

Asking the interviewer to defend their choice of algorithm or their test design reads as showing off, and it spends your turn on something you could learn in your first week. Stay with questions about how the team works.

Taking the posting's word for the role

Postings are often written from a template and list every method the team might ever touch. Until someone has described last week's work to you, you do not know which job it is.

Leaving without the timeline

Data science processes often run to several rounds, sometimes with a take-home in the middle. Close by finding out what is left and when a decision is due, so that you can fit the exercise around your other commitments.

More on this topic