Free Learning Roadmap

How to Become a DevOps / SRE

Kubernetes, Terraform, observability, on-call — keeps production up and lets engineering ship faster.

Topics
12
Resources
42
Cost
Free
Most resources are free or freemium
Format
Self-paced

1. Linux & Systems

core
Linux Fundamentals

Processes, filesystems, networking, systemd, cron, permissions. When you SSH into a broken box at 3am, this is what saves you.

Bash Scripting

Enough bash to glue things together. Sed, awk, jq, xargs, find. Not a career language — a survival language.

2. Networking

core
TCP/IP, HTTP, DNS, TLS

Every 'why can't service A reach service B' incident touches these four. Deep understanding compounds daily.

Load Balancing & Proxies

L4 vs L7, sticky sessions, health checks. nginx, HAProxy, Envoy. Cloud LB flavours — ALB, NLB, ELB, GLB.

3. Containers

core
Docker & OCI

Images, layers, multi-stage builds, image signing, distroless. The runtime everything else runs on top of.

Kubernetes

Pods, Deployments, Services, Ingress, StatefulSets, DaemonSets, HPA, PDB, network policies. The scope is huge — expect to be learning for years.

4. Infrastructure as Code

core
Terraform

Providers, modules, state, workspaces. Terraform Cloud or Spacelift for state management at scale.

Pulumi, AWS CDK

Real-programming-language IaC. Popular where teams prefer TS/Python/Go over HCL.

Helm & Kustomize

Templating and overlays for Kubernetes manifests. Every non-trivial K8s deployment uses one or both.

5. CI/CD

core
GitHub Actions

The default in most companies. Matrix builds, reusable workflows, environment protection rules.

GitOps — ArgoCD / Flux

Git is the source of truth for cluster state. Continuously reconcile from git. The modern K8s deployment pattern.

6. Cloud Platforms

core
AWS Deep Dive

IAM (deeply), VPC, EC2, EKS, RDS, S3, Route 53, CloudFront, CloudWatch, EventBridge. Multi-account org design.

GCP + Azure Basics

Multi-cloud is real for large orgs. Enough conversant literacy to work on either.

7. Observability

core
Metrics, Logs, Traces

Prometheus + Grafana, Loki (or ELK), Tempo/Jaeger. The three pillars, all with open-source implementations.

SLIs, SLOs, Error Budgets

Reliability defined by budgets, not perfectionism. This is the Google SRE thesis.

8. Incident Management & On-Call

core
On-Call Practices

Runbooks, escalation, incident commander role, blameless postmortems. Sleep depends on getting this right.

9. Security & Compliance

recommended
Cloud Security

IAM least privilege, secrets management (Vault, AWS Secrets Manager), network policies, image scanning.

Supply Chain Security

SBOMs, image signing (Cosign), SLSA framework. Increasingly compliance-mandatory.

10. Chaos & Reliability Engineering

recommended
Chaos Engineering

Deliberately inject failure — Gremlin, LitmusChaos, Chaos Mesh. Prove your reliability instead of hoping for it.

11. Cost & FinOps

recommended
Cloud Cost Management

Right-sizing, spot instances, reserved capacity, tagging, showback. Cost is a first-class ops metric.

12. Career

optional
Voices to Follow

Where SREs talk shop.

Want a personalised version?

This roadmap is the same one our platform uses internally, but the logged-in version lets you tick off topics as you complete them, track a personalised First 90 Days plan, and see your AI-durability score against this role. All free.

Open the interactive roadmap →
Curated by WhatTNext Ai · methodology · all roadmaps · last updated 2026-07-29