Data engineer interview questions, and how to answer them

Data engineering interviews test whether you can build something that keeps working after you've gone home. Expect live SQL, a pipeline design whiteboard and plenty of questions about what happens when things break. Here's what each round is looking for and how to answer the questions that come up again and again.

The process

What happens in each round

  1. 1

    Recruiter screen

    What happens

    That your stack roughly matches theirs (warehouse, orchestrator, cloud), that you've owned pipelines in production rather than just querying tables, and that pay and location line up.

  2. 2

    Live SQL or coding exercise

    What happens

    Window functions, joins that don't fan out, deduplication, handling nulls, and whether you talk through edge cases. Some teams add a Python task like parsing nested JSON from an API.

  3. 3

    Pipeline or system design

    What happens

    Whether you can take a vague business ask and design ingestion, storage, transformation and serving, with sensible choices about batch versus streaming, idempotency, backfills and monitoring.

  4. 4

    Data modeling

    What happens

    Fact and dimension tables, grain, slowly changing dimensions, and whether your model answers the questions analysts will actually ask.

  5. 5

    Behavioral and team fit

    What happens

    How you handle incidents, push back on stakeholders, document work and share on-call. They want someone who fixes the root cause instead of rerunning the job and hoping.

Questions you're likely to get

1.Walk me through a pipeline you built from source to dashboard.

Why they ask

It's the fastest way to tell whether you've owned production work or only touched one piece of it.

How to answer

  • Name the source, the volume in rough terms and how often it arrived
  • Explain how you ingested it (API pull, CDC, file drop) and where it landed raw
  • Describe the transformation layer and the tools, such as dbt models with tests
  • Say how it was scheduled and monitored, and what alerted you when it failed
  • Close with one thing that broke and what you changed afterward
2.How would you make this pipeline idempotent?

Why they ask

Pipelines get rerun constantly. They need to know a rerun won't double-count revenue.

How to answer

  • Define it plainly: running the same load twice gives the same result
  • Use merge or upsert on a stable key instead of blind inserts
  • Partition by date and overwrite whole partitions on rerun
  • Keep raw data immutable so you can always rebuild downstream
3.Write a query to return each customer's most recent order.

Why they ask

It's a classic window function test, and how you handle ties shows how careful you are.

How to answer

  • Use ROW_NUMBER partitioned by customer, ordered by order time descending
  • Filter to the first row in an outer query or with QUALIFY where supported
  • Ask how ties on the timestamp should break and add a tiebreaker column
  • Mention that RANK would return multiple rows on a tie, and when you'd want that
4.When would you choose streaming over batch?

Why they ask

Teams burn money building Kafka pipelines for reports someone reads once a morning. They want judgment, not enthusiasm.

How to answer

  • Start from the business need: how stale can the data be before it hurts
  • Streaming fits fraud checks, live inventory or user-facing features
  • Batch or micro-batch fits most reporting and is cheaper to run and debug
  • Name the operational cost of streaming: ordering, late events, state, on-call
5.Explain slowly changing dimensions and when you'd keep full history.

Why they ask

History handling is where models quietly go wrong and reports stop matching.

How to answer

  • The overwrite approach replaces the old value; you lose history but keep things simple
  • The history approach adds a new row with valid-from and valid-to dates and a current flag, so every past version survives
  • Keep history when reports must reflect the value at the time, like a customer's region when they bought
  • Mention dbt snapshots or a merge pattern as the way you'd build it
6.A dashboard shows revenue doubled overnight. How do you investigate?

Why they ask

This is the job on a bad morning. They're watching your order of operations.

How to answer

  • Check whether the pipeline ran twice or a backfill overlapped
  • Compare row counts and distinct keys at each layer to find where it jumped
  • Look for a join that fanned out after an upstream schema or key change
  • Tell stakeholders early that the number is suspect, then fix and add a test that would have caught it
7.How do you handle schema changes from a source you don't control?

Why they ask

Upstream teams rename columns without warning. It's one of the most common causes of broken loads.

How to answer

  • Land raw data in a flexible format so a new column doesn't crash ingestion
  • Add schema tests or contracts that fail loudly on renamed or dropped fields
  • Talk to the source team and agree on a heads-up process
  • Version downstream models so consumers aren't surprised
8.What makes a Spark job slow, and how do you fix it?

Why they ask

Larger shops run Spark or Databricks, and tuning is a daily chore there.

How to answer

  • Data skew, where one key holds most of the rows and one task runs forever
  • Too many small files or too few partitions
  • Wide shuffles from joins; broadcast the small side where it fits
  • Read the Spark UI stages to find the slow one before guessing
9.How do you test data pipelines?

Why they ask

Untested pipelines are the reason on-call is miserable. They want to know you'll reduce pages, not add them.

How to answer

  • Schema and constraint tests: not null, unique keys, accepted values
  • Freshness checks so a silent stall gets caught
  • Unit tests for tricky Python transforms with small fixture data
  • Reconciliation against the source for critical tables like payments
10.Tell me about a time you pushed back on a data request.

Why they ask

Stakeholders ask for real-time everything. They need someone who can say no kindly and offer something better.

How to answer

  • Set up the request and why it looked reasonable to the person asking
  • Explain the cost or risk you saw
  • Describe the alternative you proposed and how you agreed on it
  • Say how it turned out
11.How would you cut our warehouse bill?

Why they ask

Cloud warehouse spend creeps up fast, and teams want engineers who notice.

How to answer

  • Find the most expensive queries and scheduled jobs first
  • Switch full refreshes to incremental models where it's safe
  • Use partitioning or clustering so queries scan less
  • Right-size or auto-suspend compute and drop tables nobody reads

Mistakes that sink good candidates

Answering design questions with a list of tools instead of reasoning from what the business needs

Claiming you've never had a pipeline fail, which tells them you haven't run one in production

Writing SQL in silence during the live round instead of saying what you're checking for

Badmouthing the analysts or source teams you worked with

Need more Data Engineer interviews to prep for?

HeroApply applies to Data Engineer jobs that match you, every day.

Find Data Engineer jobs