Site Reliability Engineer
About this role
Employer-provided description, formatted for easier reading.
About The Role
The Site Reliability Engineer owns the uptime, performance, and scalability of production systems handling thousands of requests per second. The role blends software engineering with systems thinking - eliminating toil through automation, designing infrastructure that fails gracefully, and running incident response that keeps customers informed and services recovering fast.
This is a hands-on role on a platform team that other engineering teams depend on. You will shape how services are deployed, monitored, and scaled across Kubernetes and cloud infrastructure, and your work directly determines whether engineers can ship safely and customers stay online.
Key Responsibilities
- Build and maintain scalable infrastructure on AWS/GCP using Terraform, treating everything as code with peer-reviewed modules and CI-driven provisioning
- Own SLOs and error budgets for critical services - define SLIs, build dashboards in Grafana/Datadog, and drive reliability decisions across product teams
- Design and operate Kubernetes clusters: autoscaling, resource tuning, pod disruption budgets, and Helm-based service deployment standards
- Lead incident response as an on-call escalation point - run incident command, write blameless postmortems, and drive remediation work to closure
- Automate toil away with Python, Bash, and Go - self-healing runbooks, capacity management, and cost optimization tooling
- Harden production systems: implement network policies, secrets management, least-privilege IAM, and vulnerability remediation pipelines
- Partner with development teams on production readiness - review designs for reliability, set deployment standards, and improve observability coverage
What We Are Looking For
- 3–7 years of experience in SRE, DevOps, or infrastructure engineering, including operating large-scale production systems with real on-call responsibilities
- Deep hands-on experience with Kubernetes in production - operating clusters, debugging workloads, and tuning for performance and cost
- Strong infrastructure-as-code skills with Terraform (or similar) and experience with CI/CD pipelines (GitHub Actions, GitLab CI, or ArgoCD)
- Solid grasp of distributed systems fundamentals: load balancing, service meshes, queues, caching strategies, and failure modes
- Production experience with observability stacks: Prometheus, Grafana, Datadog, or equivalent - including building actionable alerts, not noisy ones
- BS in Computer Science or equivalent practical experience; scripting proficiency in Python, Go, or Bash
- Bonus: experience with chaos engineering, multi-region architectures, service mesh (Istio/Linkerd), or database reliability (PostgreSQL, MySQL)