DevOps Engineer interview questions and how to answer them

DevOps interviews are less about reciting tool features and more about proving you've been close to production when it broke. Expect a mix of whiteboard design, live troubleshooting and stories about incidents you were part of. Here's what each round is really checking, and how to answer the questions that come up again and again.

The process

What happens in each round

  1. 1

    Recruiter screen

    What happens

    Whether your stack matches theirs closely enough. They'll ask which cloud you've used, whether you've run Kubernetes in production or only in a tutorial, and how you feel about on-call.

  2. 2

    Hiring manager conversation

    What happens

    How you think about reliability versus speed, and whether you can explain a past project without hiding behind the team. They want to hear what you personally built or changed.

  3. 3

    Technical troubleshooting round

    What happens

    How you debug under pressure. Often a broken container, a failing pipeline or a Linux box that's out of memory, shared in a terminal or described out loud. Your order of checks matters more than the final answer.

  4. 4

    System or pipeline design

    What happens

    Whether you can design a CI/CD flow or an environment layout from scratch, name the tradeoffs, and plan for rollback and secrets from the start.

  5. 5

    Team or values interview

    What happens

    How you handle blame, pushback from developers and incidents that happen on your watch. Blameless postmortem habits carry a lot of weight here.

Questions you're likely to get

1.Walk me through the CI/CD pipeline you know best, from commit to production.

Why they ask

It's the fastest way to see whether you've built pipelines or just watched them run.

How to answer

  • Name the tool and the trigger, for example a pull request in GitHub Actions or GitLab CI
  • Describe the stages in order: lint, unit tests, image build, security scan, deploy to staging, promotion
  • Say how artifacts are versioned and stored so production runs the exact image that was tested
  • Explain how a release gets rolled back and who can approve a production deploy
  • Mention one thing you'd change about it and why
2.A deploy to Kubernetes is stuck and pods keep restarting. What do you check, in what order?

Why they ask

This is the classic troubleshooting prompt. They're listening for a method, not a lucky guess.

How to answer

  • Start with kubectl get pods and kubectl describe pod to read the events and the restart reason
  • Check the container logs, including the previous container's logs after a crash
  • Look at probes, resource limits and whether the pod was killed for running out of memory
  • Confirm config: missing secrets, wrong environment variables, a bad image tag
  • Decide early whether to roll back while you keep digging
3.How do you manage Terraform state on a team?

Why they ask

State mistakes cause real outages and lost resources. People who've been burned answer differently.

How to answer

  • Store state remotely, for example in a cloud bucket with locking, never on a laptop
  • Split state by environment or component so a bad apply has a smaller blast radius
  • Run plan in CI on every pull request and have a human read the diff before apply
  • Explain how you'd handle drift or import a resource someone created by hand
4.How do you handle secrets in pipelines and running services?

Why they ask

Leaked credentials in a repo or a build log are among the most common self-inflicted security incidents.

How to answer

  • Keep secrets out of Git entirely and scan for them in CI
  • Use a secrets manager such as HashiCorp Vault or the cloud provider's own service
  • Prefer short-lived credentials and workload identity over long-lived keys
  • Explain how you'd rotate a secret that leaked, step by step
5.Tell me about an incident you were part of. What happened, and what changed afterward?

Why they ask

They want proof you've been in the room when things broke, and that you learn from it without pointing fingers.

How to answer

  • Set the scene briefly: what service, what users saw
  • Say what you personally did, including the first decision you made
  • Explain the root cause in plain terms
  • Name the follow-up work, such as a new alert, a canary step or a runbook, and whether it got done
6.How would you design monitoring and alerting for a new service?

Why they ask

Noisy alerts burn out on-call teams. They want someone who alerts on symptoms users feel.

How to answer

  • Start from what users care about: error rate, latency, availability
  • Collect metrics with something like Prometheus and put them on Grafana dashboards
  • Page only on alerts that need a human now; send the rest to a ticket queue
  • Attach a runbook link to every page so the person woken up knows where to start
7.A developer says the pipeline is too slow and they're thinking of skipping tests. What do you do?

Why they ask

DevOps is a service to developers. They're checking whether you fix the pain or just enforce rules.

How to answer

  • Take the complaint seriously and measure where the time goes
  • Look at caching dependencies, parallel test jobs and smaller Docker images
  • Split fast checks that block merges from slow suites that run after
  • Explain why skipping tests just moves the cost to production
8.What's the difference between blue-green and canary deployments, and when would you pick each?

Why they ask

It checks whether you understand release risk, not just release mechanics.

How to answer

  • Blue-green switches all traffic between two full environments, which makes rollback a quick switch back
  • Canary sends a small slice of traffic to the new version first and widens it if metrics hold
  • Blue-green costs more infrastructure; canary needs good metrics to judge the slice
  • Mention database changes as the hard part in both
9.A Linux server is slow and nobody knows why. Where do you start?

Why they ask

Plenty of DevOps candidates know Kubernetes but not the box underneath it.

How to answer

  • Check load, memory and swap with top or htop
  • Look at disk space and disk wait with df and iostat
  • Read recent logs with journalctl and check for runaway processes
  • Ask what changed recently before assuming it's hardware
10.How do you decide what to automate next?

Why they ask

Automation eats time too. They want someone who picks the work with the best payoff.

How to answer

  • Look for tasks that are frequent, error-prone or blocking other people
  • Weigh the build and upkeep cost against the time saved
  • Give an example of something you chose not to automate, and why
11.Tell me about a time you disagreed with a developer or a manager about a release.

Why they ask

You'll sometimes be the person saying not yet. They want to know you can do it without making enemies.

How to answer

  • Describe the risk you saw and the evidence behind it
  • Explain how you offered a path forward, such as a feature flag or a staged rollout
  • Say how it ended, even if you lost the argument

Mistakes that sink good candidates

Talking about tools you've only used in a tutorial as if you've run them in production

Blaming a developer or another team in your incident story

Skipping rollback and secrets in a design answer

Saying you don't mind on-call, then being unable to describe a single page you handled

Need more DevOps Engineer interviews to prep for?

HeroApply applies to DevOps Engineer jobs that match you, every day.

Find DevOps Engineer jobs