Site Reliability Engineer
Spotted 3d agoFull-time
Job details
- Employment
- Full-time
- Level
- Mid level
- Experience
- 3+ years
- Education
- Bachelor's degree
- Posted
- Oct 7, 2026
- Last confirmed open
- Oct 8, 2026
Job description
About this role
About The Role
The role focuses on keeping large-scale production systems reliable, observable, and fast. This team owns the infrastructure layer - Kubernetes clusters, CI/CD pipelines, and the observability stack that hundreds of internal engineers depend on to ship daily.
This is a hands-on position for someone who treats reliability as a product problem: defining SLOs, instrumenting everything, automating away toil, and taking pages when systems misbehave. The team sits directly on-call for services handling significant production traffic.
Key Responsibilities
- Operate and scale Kubernetes-based infrastructure across multiple environments, including cluster upgrades, autoscaling policies, and node lifecycle management
- Build and maintain Terraform modules for provisioning cloud infrastructure (AWS or GCP), enforcing infrastructure-as-code practices across the org
- Design SLOs, SLIs, and error budgets with product teams, and drive reliability improvements through error budget policy decisions
- Instrument services with Prometheus, Grafana, and OpenTelemetry; build dashboards and alerts that page on symptoms, not causes
- Lead incident response for production outages, write blameless postmortems, and follow through on action items to closure
- Automate operational toil out of existence - runbooks, self-service tooling, and pipelines that remove manual steps from deploys and scaling operations
- Improve CI/CD pipelines (GitHub Actions, ArgoCD) to make deployments faster, safer, and progressively rolled out by default
What We Are Looking For
- 3–7 years of experience in SRE, DevOps, or infrastructure engineering, including ownership of production systems with real on-call rotations
- Deep hands-on experience with Kubernetes in production: operating clusters, troubleshooting workloads, and managing controllers/operators
- Strong infrastructure-as-code skills with Terraform (or equivalent), plus scripting proficiency in Python, Go, or Bash
- Experience running observability tooling in production: Prometheus, Grafana, distributed tracing (OpenTelemetry or similar)
- Practical understanding of Linux systems, networking fundamentals (DNS, TLS, load balancing), and container internals
- Track record of leading incident response and writing postmortems that actually change engineering practice
- Bonus: Experience with service meshes (Istio/Linkerd), chaos engineering, GitOps workflows, or cloud cost optimization; BS in Computer Science or equivalent practical experience
Interested in this role?Continue on LinkedIn to apply.
Apply on LinkedIn