Questions to Ask in a Data Engineer Interview
For candidates in a data engineering interview loop who have reached the part where the interviewer asks what you want to know. The groups follow the people you meet: the role and its backlog for the hiring manager, then pipelines, data quality and on-call for the engineers who run them, then the analysts and data scientists who depend on the output, and last the level, growth and remaining rounds, where the recruiter comes in. Two or three per conversation is plenty, and the on-call group is the one to finish before you say yes.
Want questions from the whole vault instead? Try the random question generator.
The questions
Each question, and why to ask it
The role
What would a data engineer on this team spend most of a normal week doing?
Why ask it
Ask about last week in particular, since an average week is easy to tidy up. Sort what you hear into building pipelines, fixing ones that broke and answering requests for data, because the title covers all three in very different proportions. If the manager and an engineer on the loop give different splits, the engineer's is the one you would live.
What changed that made this the moment to hire another data engineer?
Why ask it
More data, a migration, someone leaving and a new product line each put you on different work. A departure is worth one follow-up about what that person owned, since those pipelines are likely to be yours within the month. If this is the first data engineering hire, ask who has been keeping the loads running until now.
Is this role closer to platform work, pipelines for specific teams, or analytics engineering?
Why ask it
Platform work is orchestration, infrastructure and tooling that other engineers build on. Pipeline work is sources, loads and deadlines. Analytics engineering is modeling tables for the people who query them, mostly in SQL. Most roles mix the three, so ask which one the manager would protect in a week that ran short.
What would be the first pipeline or project with my name on it?
Why ask it
A named project with a consumer waiting for it means the role has been thought through. 'You will pick things up from the backlog' is workable too, but ask who chooses, and whether the first pick is something a newcomer can finish without knowing the whole warehouse.
Six months from my start date, what should be in production that is not there today?
Why ask it
This hands you the manager's yardstick before you start. Good answers are things you could point at: a source brought in, a migration finished, a job that no longer fails every week. If the answer is all about learning the systems, ask what comes after the learning.
How many data engineers are on the team, and roughly how many pipelines and sources do they keep running between them?
Why ask it
The ratio is what you are after. A handful of engineers looking after hundreds of jobs is a team that runs things, whatever the posting says about building. Ask who else writes pipelines besides this team, since analysts and backend engineers sometimes do, and who gets the alert when one of theirs fails.
What would my first month look like, and how long before a new engineer usually ships a change to a production pipeline?
Why ask it
Access is often what holds a new data engineer up: warehouse roles, cloud accounts, credentials for each source. A team that can say 'first change in week two' has set those up ahead of time. If the last hire waited a month for permissions, ask what has been fixed since.
What is at the top of the backlog right now, and who decides what comes off it next?
Why ask it
The top items are probably your first quarter. Who puts them in order matters as much: a manager with a roadmap, or whichever stakeholder asked most recently. Ask how often something jumps the queue, and who is allowed to make it jump.
Which piece of tech debt costs the team the most time each week?
Why ask it
Put it to the manager and to an engineer if you can, because the engineer's answer tends to come faster and with more detail. A specific answer, such as a nightly job written years ago that nobody dares to change, means the team knows where it hurts. Then ask whether retiring it is on any plan, and when a piece of debt was last paid off on purpose.
Pipelines
Can you walk me through the stack from source to dashboard: ingestion, orchestration, warehouse and transformation?
Why ask it
Let them draw it if there is a whiteboard or a shared screen. Count how many tools do the same job, which usually marks where an old approach was never fully replaced. Notice which layer the interviewer hurries past, and come back to it.
Which languages is the work written in day to day: SQL, Python, Scala, or something else?
Why ask it
A posting tends to list every language the team has ever touched, and this gets you the proportions. A job that is nine parts SQL inside a transformation tool is not the same job as one writing Spark code or services, and the interview rounds do not always tell you which it is. If one language on the list is fading, ask what still runs on it.
How much of the pipeline work is batch, and how much is streaming?
Why ask it
The two call for different skills and fail in different ways. If streaming is on the posting, ask what runs on it today and who needs data that fresh. A stream that feeds a report someone reads once a day deserves a gentle question about why.
Is a migration under way, and how far along is it?
Why ask it
A move between warehouses, orchestrators or clouds can be the most interesting work on offer, or a long stretch of running two systems side by side. Ask when it started, what the first finish date was and what is left. The distance between that date and today shows how the team estimates.
Roughly how much data moves through in a day, and which job is closest to its limit?
Why ask it
Scale decides whether the hard problems are about performance or about correctness and people. The second half is the useful half: a job that barely finishes before the morning deadline is where the next incident comes from, and possibly your first project.
How does a pipeline change get from my laptop to production, and is there somewhere to try it against realistic data first?
Why ask it
Have the engineer describe the last change they shipped themselves. Version control, review by a second person, automated tests and a staging copy with real volumes are the things to listen for. Where testing means running it in production and watching, mistakes are found by the people reading the dashboards.
When a job fails halfway, can it be rerun safely, and how is a backfill done?
Why ask it
It is a quiet test of how the pipelines were designed. Reruns that are safe and backfills that take one command mean someone planned for failure. If the answer involves deleting rows by hand first, ask how often that has to happen and who is trusted to do it.
How are tables modeled in the warehouse, and who decides when a new one gets added?
Why ask it
Listen for a named approach and a review step, or an admission that anyone can create anything. A warehouse with no gatekeeping fills up with near-copies of one table, and working out which copy is right becomes part of every task.
If I needed to know where a column comes from, where would I look?
Why ask it
A lineage tool, a catalog, readable transformation code or 'ask whoever built it' are the possible answers. The last is common and survivable, but find out whether that person still works there. Offering to document what you trace in your first months tends to go down well.
Who watches the warehouse and cloud bill, and has cost ever changed a design?
Why ask it
Where engineers can see what their own queries and jobs cost, efficiency is part of the craft. Where nobody looks until finance complains, expect a sudden cleanup project at some point. One example of a design reworked to save money tells you which kind of team this is.
Which pipelines carry personal or regulated data, and who signs off before a new source like that is loaded?
Why ask it
What is required depends on the country, the industry and the company's own policies, so ask how it is done there instead of assuming. The answer you are hoping for has a named owner and a routine, such as masking outside production or a review before the first load. With neither, those judgments fall to whoever writes the pipeline.
Data quality
Who owns data quality here: the teams that produce the data, this team, or whoever notices first?
Why ask it
Few answers change the job as much as this one. Where the producing teams share ownership, problems get fixed at the source. Where anything wrong in any table is this team's fault by default, you would be cleaning up after systems you have no power to change.
What was the last wrong number to reach a dashboard, and how long was it there before someone caught it?
Why ask it
A team that can tell this story with dates has looked hard at its own failures. Listen for who caught it. An automated check is the good answer, an analyst is the usual one, and an executive in a meeting is the one that gets budgets approved for testing.
Which checks run on data before it is published: freshness, row counts, schema, business rules?
Why ask it
Each of the four catches a different failure, so ask which exist and which stop the load when they fail. Checks that only warn tend to be muted after a few noisy weeks. Also ask who can add a check, and whether analysts are allowed to.
When an upstream team changes a schema, does this team hear before or after something breaks?
Why ask it
Before means there is a contract, a review or at least a habit of telling people. After means data engineers find out from a failed job, and a fair part of the role is detective work. Ask what happened the last time, and whether the upstream team changed anything because of it.
Do any datasets carry a promise about freshness or accuracy, and who agreed to it?
Why ask it
A written promise, such as finance tables ready by a set hour, shows which pipelines are treated as critical and which may slip. If the promises exist only in stakeholders' heads, you would learn them one complaint at a time, so ask which three tables cause the loudest messages when they are late.
Once a table is found to be wrong, how do the people who already used it find out?
Why ask it
Repairing the pipeline is half of a data incident. The other half is the reports, exports and decisions built on the bad rows. A team with a way of marking affected tables and telling their readers has been through this before and kept the lesson.
Where do metric definitions live: in the transformation code, a semantic layer, or each dashboard?
Why ask it
One shared place means a definition is changed once and everybody gets it. Definitions scattered across dashboards let one metric come out two ways, and data engineers get called in to referee. If it is scattered, ask whether pulling it together is planned and whose project that would be.
On-call
Who carries the pager for the pipelines, and how often would my turn come around?
Why ask it
The headcount sets how often it is you, so do the arithmetic while they talk. Ask whether a turn is a day or a week and whether weekends are part of it. With only two people sharing it, you would be on call half the year.
How soon would a new hire join the rotation, and what comes before the first shift?
Why ask it
Shadowing a shift or two before carrying the pager alone means someone has thought about it. Ask whether runbooks exist for the common failures and when one was last edited. Being on call in the second week with no notes says something about more than on-call.
How many alerts fired last week, and how many of them needed a person to do something?
Why ask it
Put it to an engineer who is on the rotation, who will know without looking it up. A wide gap between the two numbers means noisy alerting, and tired people miss the one alert that counts. Then ask whether anyone is given time to tune the alerts.
If a nightly load fails at 3 a.m., does someone fix it then, or does it wait for the morning?
Why ask it
It depends on who is waiting for the data and by when, so a good answer names a consumer and an hour. Teams that have agreed which jobs can wait sleep better than teams that treat every failure as urgent. Ask which jobs are on the wake-up list and how they got there.
Does on-call cover only this team's pipelines, or the source systems and dashboards as well?
Why ask it
Scope is a common on-call grievance in data work, because a broken dashboard looks the same whoever broke it. Ask where a problem is handed off when the cause is upstream, and whether the other team has anybody awake to take it.
After an incident, is there a write-up, and does the fix make it onto the schedule?
Why ask it
Have them describe the most recent one. The document matters less than whether the follow-up work was done, so ask what is different now. One job failing three times for one reason is an answer about priorities, whatever the process says on paper.
Is on-call paid on top of salary here, and do people get time back after a rough night?
Why ask it
Employers handle this differently and local rules can come into it, so take what you hear as this company's practice and ask to see it in writing alongside any offer. The recruiter usually knows the policy. An engineer knows whether people really take the time back.
Consumers
Who uses what this team builds: analysts, data scientists, product engineers, finance?
Why ask it
Each group wants something different from a data engineer. Analysts want modeled tables, data scientists want raw history and features, and finance wants numbers that never move after the books close. Whichever group the interviewer names first is usually the one that sets the team's priorities.
How do requests from analysts and data scientists reach the team: tickets, a channel, or a direct message to whoever they know?
Why ask it
Where requests arrive by direct message, the newest engineer gets none at first and then, once known, far too many. A queue with somebody triaging it protects your uninterrupted hours. Ask how long a typical request for a new table sits before work begins.
Where does this team hand off to analytics: at raw tables, at modeled tables, or at the dashboard?
Why ask it
Some teams land the raw data and stop. Others own everything up to the chart, which is a different job with more SQL and more time with stakeholders. Ask where the line fell on a real project from the past few months, since the org chart and the practice can differ.
Can analysts and data scientists build their own transformations, or does everything go through this team?
Why ask it
A self-serve setup turns data engineers into platform builders and reviewers. A gatekept one turns them into a service desk with a queue. If it is self-serve, ask who cleans up when somebody's query slows the warehouse for everyone on a Monday morning.
When a data scientist's model needs fresh data every day, who builds and runs that pipeline?
Why ask it
Machine learning brings feature pipelines, snapshots of training data and jobs that fail in ways a reporting load does not. Find out whether that work belongs to data engineering, to a separate machine learning engineering group or to the data scientist. If it would be yours, ask how many such pipelines run today.
When two dashboards disagree, who gets asked to explain it, and how much of the week do requests like that take?
Why ask it
A little of this is healthy, because it keeps engineers close to what the data is for. When it swallows most of the week, planned work slips and nobody can say why. Ask whether one person takes these each week so the others can concentrate.
Are data engineers in the room when a product feature is designed, or do they learn about new data at launch?
Why ask it
Ask for the last feature where data engineering was consulted before the build, and what that changed in the design. A team that only hears at launch spends its time reverse-engineering another team's logging against a deadline, and the events rarely arrive in a shape that can be used as they are.
Does the data team sit under engineering, analytics or somewhere else, and who does its head report to?
Why ask it
A data team inside engineering more often treats pipelines as software, with the review and tooling that implies. One inside analytics or finance sits nearer the people it serves and sometimes further from those habits. Neither placement settles anything alone, so ask who argued for the last tool or hire the team got, and whether they won.
Growth and next steps
How is a data engineer's performance judged here, when a good month is one where nothing broke?
Why ask it
Reliability is invisible while it works, so ask what the manager points to at review time: pipelines delivered, incidents avoided, what consumers say. Review systems differ from one employer to the next, so ask how the last cycle ran on this team. If only launches count, maintenance becomes the work nobody volunteers for.
What does a senior data engineer here do that a mid-level one does not?
Why ask it
A real answer is about scope: owning a domain, designing across systems, setting standards other people follow. If it is mostly years served, promotion may be a matter of waiting. Ask what the last person promoted to senior had built or taken over.
Who is the most senior data engineer here who does not manage people?
Why ask it
A name and a title mean the technical track exists in practice. Ask what that person works on and whether you would work alongside them. If every senior engineer became a manager, management is the only path the company has actually built.
Who on the team would push back on a pipeline design of mine before it gets built?
Why ask it
On a small data team you can be the most experienced person on day one, which is flattering and lonely. If nobody would push back, ask where the team gets a second opinion. If there is a name, try to meet that person before you decide.
Which level is this opening, and what is the pay range for it?
Why ask it
This belongs with the recruiter, and early in the process. Whether a range has to be shared depends on where the job is, so if none is offered, ask for the level and work from there. Ask too whether the figure is base salary alone, because two ranges that look alike can be counting different things.
Which rounds are still ahead, and do they include SQL, data modeling or a pipeline design exercise?
Why ask it
Another one for the recruiter. Ask how long each round runs, whether the design exercise is on a whiteboard or in a shared document, and when a decision is expected after the last one. Knowing the format lets you practice the right thing instead of everything.
Which part of the data platform surprised you most after you joined?
Why ask it
Save it for an engineer, near the end, when the formal part is over. A surprise measures the distance between what the interviews described and what the job turned out to be. A pleasant one is worth as much to you as an unpleasant one.
From what you have heard today, where would I have the most to learn on this stack?
Why ask it
It brings a doubt into the open while you can still answer it. Reply with one concrete example, or say plainly how you would get up to speed, and leave it there. Whatever they name is also what to read up on before the next round.
Getting straight answers about a data engineering job
Practical guidance for the conversation itself
Matching each group to an interviewer
Hiring manager: scope, backlog and the yardstick
The manager knows why the seat exists, what is queued up for it and what a good first year would look like. The role group and the performance questions in Growth and next steps are theirs. They can describe the pipelines too, but usually as they are meant to work.
Engineers on the loop: pipelines, quality and the pager
The people who ship the jobs and carry the alerts know how things really run. Give them Pipelines, Data quality and On-call, and ask about their own last change, their own last alert. If the loop includes an analyst or a data scientist, the Consumers group works in reverse: ask how long their last request to data engineering took.
Recruiter: level, pay band and rounds
A recruiter can rarely say how backfills are done, and should not be asked. They can tell you the level, the range, what the remaining rounds test and what the written on-call policy is. Settle those by phone or email so that your time with the engineers is spent on engineering.
Cross off what the posting already told you
A posting that names an orchestrator, a warehouse and a transformation tool has answered the plain stack question. Ask the next one instead: which of those is being replaced, and which one the team wishes it had never adopted.
Lead with the answer that could end it
A forty-five minute technical round often leaves five minutes for you. Go in with three questions for that interviewer and ask first the one whose answer could make you turn the job down. For many data engineers that is who owns data quality or how heavy the pager is.
How to ask about pipelines and incidents
Ask for the last failure, not the architecture
An architecture description is the system on its best day. 'What broke most recently, and how did you find out?' gets you the monitoring, the ownership and the mood of the team in one story. Most questions in Data quality and On-call work best in that form.
Use the design round as a way in
If you were set a pipeline design exercise, ask afterwards how close it was to something the team runs. Interviewers enjoy the comparison, and you learn where their real system departs from the tidy answer you were steered toward.
Ask to see last night's run
On a video call, ask whether the engineer can show the scheduler's view of last night's jobs. Some companies do not share a screen with candidates, and that is a fair answer. Where they do, a page of green with two red retries on it tells you more about the alerts, the reruns and the mood than ten minutes of description.
Put the ownership question to both sides
Ask the manager who owns data quality, then ask an engineer. A manager may describe the agreement with the producing teams, and an engineer may describe what happens when a load fails on a Sunday. Where the two accounts part ways is the thing to raise before accepting.
Say what you have already lived with
'I have carried a pager for nightly loads and would do it again. How heavy is yours?' gets a franker answer than the bare question. The interviewer stops wondering whether you are hunting for a reason to say no and starts describing the rotation. Say it only if it is true.
What the answers add up to
Count the names
After the interviews, list source data, pipelines, quality checks and metric definitions, and write beside each the person or team named as responsible. Messy pipelines with an owner get better. Blanks on that list are the parts of the job that would quietly become yours.
Decide whether the job is build, run or serve
Put the role answers and the consumer answers side by side and decide which word fits: building new pipelines and platform, running what exists, or serving requests from other teams. All three are honest work. Trouble comes from accepting one while picturing another.
Turn the on-call answers into nights
Rotation size, alerts that needed a person and the wake-up list let you estimate interrupted nights in a month. Work the number out before an offer arrives, and compare it with what the company gives back for them.
Check the stack against where you want to be
A team in the middle of a migration can teach you two systems and the judgment to move between them. A settled stack offers depth in one. Decide which you want on your resume in three years, then look again at the migration and streaming answers.
Ways to waste your turn
Stopping at the tool list
Knowing which orchestrator a team runs says little about what the job is like. Two teams on identical tools can differ completely in who owns quality and who gets woken up. Spend one question on the stack and the others on how it is operated.
Arguing with their choices
Asking why they have not moved to a tool you prefer turns your questions into a critique of people you hope to work with. 'What would you choose differently if you were starting today?' gets the same information and leaves them the ones doing the criticizing.
Leaving on-call until after the offer
Candidates skip it for fear of seeming unwilling. Asked plainly of an engineer, it reads as experience: people who have carried a pager for data pipelines know to ask. Learning the rotation in your first week is too late to weigh it.
Taking the diagram for the system
A clean picture of sources, warehouse and dashboards is how the platform is supposed to work. The quality and incident answers tell you how it does work. When the two disagree, go with the incidents.