Cloud Engineer interview questions and how to answer them

Cloud engineer interviews test whether you can reason about systems you can't see, under pressure, with money and security on the line. Expect a mix of whiteboard design, hands-on Terraform or Linux work, and a lot of "tell me about a time it broke." The questions below are the ones that come up again and again, with what the interviewer is really listening for.

The process

What happens in each round

  1. 1

    Recruiter screen

    What happens

    Which cloud you've worked on, whether you've written infrastructure as code or only clicked in the console, how you feel about on-call, and whether your pay expectations fit the band.

  2. 2

    Hiring manager conversation

    What happens

    The real scope of your past work. They'll push on one project until they find the edge of what you personally did versus what your team did.

  3. 3

    Technical round or live exercise

    What happens

    Networking, identity and troubleshooting. Often you'll debug a broken Terraform plan, read a failing pipeline, or explain why a service can't reach its database.

  4. 4

    System design

    What happens

    Whether you can lay out a highly available setup for a web app, talk through tradeoffs in cost and failure, and say what you'd monitor.

  5. 5

    Team or behavioral round

    What happens

    How you handle incidents, push back on developers without blocking them, and write things down so the next person isn't guessing.

Questions you're likely to get

1.Walk me through how you'd design the network for a new web application in AWS.

Why they ask

VPC design is the foundation everything else sits on. A shaky answer here tells them you've only worked inside networks someone else built.

How to answer

  • Start with one VPC split across several availability zones so a single zone failure doesn't take you down
  • Public subnets for the load balancer only, private subnets for app servers and databases
  • NAT gateway for outbound traffic from private subnets, and a note on what it costs
  • Security groups scoped tightly, with the database accepting traffic only from the app tier
  • Leave room in the address range for future growth and peering
2.A developer says their app can't reach the database. How do you troubleshoot it?

Why they ask

This happens constantly. They want to see a calm, ordered process instead of random clicking.

How to answer

  • Confirm the exact error: timeout usually means network, refused or auth errors mean something else
  • Check DNS resolution from the app host
  • Walk the path: security groups, network ACLs, route tables, subnet placement
  • Test from the host with a simple connection tool before touching config
  • Check database-side limits, credentials and whether it's even running
3.How do you manage Terraform state on a team, and what goes wrong when you don't?

Why they ask

Anyone can write a Terraform resource. Running it safely with other people is the real skill.

How to answer

  • Remote state in a locked backend, such as an encrypted bucket with a lock table or Terraform Cloud
  • State locking so two people can't apply at once
  • Splitting state by environment and by component so a mistake has a small blast radius
  • Plans reviewed in pull requests and applied from the pipeline, not laptops
  • A story about drift or a corrupted state and how you recovered
4.How would you give an application access to a storage bucket without putting keys in the code?

Why they ask

Leaked long-lived keys cause a lot of real breaches. They want to hear that you reach for roles by default.

How to answer

  • Attach an IAM role to the compute, whether that's an instance profile, a task role or a workload identity in Kubernetes
  • Grant least privilege: the one bucket, the actions it needs
  • Use short-lived credentials and OIDC for pipelines instead of stored secrets
  • Keep anything that must be secret in a secrets manager with rotation
5.Our cloud bill went up sharply last month. Where do you start?

Why they ask

Cost is part of the job whether the posting says so or not. This shows whether you think like an owner.

How to answer

  • Open the cost explorer and break the jump down by service, account and tag
  • Look for the usual suspects: idle instances, unattached volumes, data transfer, NAT traffic, forgotten snapshots
  • Fix tagging so the next spike points at a team
  • Suggest guardrails like budgets and alerts, and rightsizing or reserved capacity for steady workloads
6.Tell me about an outage you were part of. What happened and what changed afterward?

Why they ask

Everyone who has run production has a story. They're checking your honesty and whether you learn from it.

How to answer

  • A short, clear timeline: what alerted, what you saw first, what you tried
  • Your part specifically, including anything you got wrong
  • How you communicated with the people affected while it was happening
  • The root cause and the lasting fix, like a new alert, a runbook or a guardrail in the pipeline
  • Mention a blameless postmortem if your team ran one
7.How would you make a web application highly available across failures?

Why they ask

This is the classic design prompt. They want tradeoffs, not a list of services.

How to answer

  • Stateless app servers in an auto scaling group or Kubernetes deployment across zones
  • A managed database with a standby in another zone and tested failover
  • Health checks on the load balancer that actually reflect whether the app works
  • Say when multi-region is worth it and when it's expensive overkill
  • Backups that you've restored at least once, not just taken
8.When would you choose Kubernetes, and when would you avoid it?

Why they ask

Teams have been burned by Kubernetes they didn't need. Judgment matters more than enthusiasm.

How to answer

  • Good fit: many services, a platform team to run it, need for consistent deploys and scaling
  • Poor fit: a handful of apps and a small team, where managed containers or serverless is simpler
  • Name the ongoing cost: upgrades, networking, observability, security patching
  • Mention a managed option like EKS, AKS or GKE if they do go that way
9.What does a good CI/CD pipeline for infrastructure look like to you?

Why they ask

They want to know if changes reach production through review or through someone's laptop.

How to answer

  • Formatting, validation and a policy or security scan on every pull request
  • Terraform plan posted to the pull request for review
  • Apply only from the main branch, with approval for production
  • Pipeline authenticates with short-lived credentials, not stored keys
10.A developer wants admin access to production to move faster. How do you handle it?

Why they ask

You'll be in the middle of security and speed every week. They want to see you protect the account without becoming the bottleneck.

How to answer

  • Find out what they're actually trying to do
  • Offer a scoped role or a self-service path that covers that need
  • Use temporary, logged elevation for break-glass situations
  • Explain the risk in terms of their own service, not policy
11.How do you decide what to alert on?

Why they ask

Noisy alerts burn out on-call engineers. This is where they see whether you've carried a pager.

How to answer

  • Page on symptoms users feel, like error rate and latency, not every CPU spike
  • Send lower-urgency signals to a channel or ticket instead of a phone
  • Every page should have a runbook and a clear action
  • Review alerts after incidents and delete the ones nobody acts on
12.Tell me about something you automated that people used to do by hand.

Why they ask

Automation is the heart of the job. They want a concrete example with a before and after.

How to answer

  • What the manual task was and why it hurt, such as errors, delays or late-night work
  • The tool you picked, like Terraform, Ansible, Python or a pipeline, and why
  • How you rolled it out so people trusted it
  • The result in plain terms: fewer mistakes, faster turnaround, one less on-call page

Mistakes that sink good candidates

Answering design questions with a list of service names and no reasoning about failure or cost

Taking credit for a whole migration when you owned one piece of it; they'll find out in follow-up questions

Dismissing security steps as slowing things down

Freezing in the live exercise instead of talking through what you'd check next

Need more Cloud Engineer interviews to prep for?

HeroApply applies to Cloud Engineer jobs that match you, every day. 1,629 jobs are open today.

Find Cloud Engineer jobs