Site Reliability Engineer
Spotted 13m agoFull-time
Job details
- Employment
- Full-time
- Level
- Mid level
- Experience
- 3+ years
- Education
- Bachelor's degree
- Posted
- Oct 11, 2026
- Last confirmed open
- Oct 11, 2026
Job description
About this role
About The Role
The role owns the reliability of large-scale production systems — distributed services, Kubernetes-based infrastructure, and the pipelines that deploy, observe, and heal them. SLOs, error budgets, and incident response are not buzzwords here; they are the operating model.
You will join a platform reliability team that partners directly with product engineering groups, ensuring that latency, availability, and infrastructure cost targets hold under real traffic — millions of requests per day — while the platform continues to ship.
Key Responsibilities
- Design, implement, and maintain SLOs and error budgets for critical services, driving prioritization of reliability work across engineering teams
- Build and evolve observability tooling — Prometheus, Grafana, OpenTelemetry, distributed tracing — so on-call engineers can diagnose issues in minutes, not hours
- Automate infrastructure provisioning and configuration with Terraform, Ansible, and GitOps workflows on Kubernetes (EKS/GKE)
- Lead and participate in incident response: drive mitigations, write blameless postmortems, and track remediation items to completion
- Improve deployment pipelines — CI/CD with GitHub Actions or GitLab CI, canary and blue-green strategies, automated rollback — to reduce deployment risk without slowing release cadence
- Perform capacity planning, load testing, and performance tuning for stateful and stateless services across multi-region cloud environments (AWS or GCP)
- Author runbooks and automation that eliminate toil, progressively shifting pager load from humans to self-healing systems
What We Are Looking For
- 3–6 years of experience in SRE, DevOps, or backend infrastructure engineering, including running production systems at scale with on-call responsibility
- Strong proficiency with Kubernetes in production: cluster operations, Helm, autoscaling, networking, and troubleshooting under pressure
- Deep hands-on experience with at least one major cloud provider (AWS, GCP, or Azure), including IAM, VPC networking, and managed data services
- Solid coding skills in Python, Go, or Bash for building tooling, automation, and infrastructure-as-code (Terraform required)
- Proven experience with monitoring and observability stacks (Prometheus/Grafana, Datadog, or equivalent) and defining meaningful service-level metrics
- Track record of leading incident response and writing clear, actionable postmortems; strong written communication is essential
- Bachelor's degree in Computer Science, Engineering, or equivalent practical experience. Bonus: experience with chaos engineering, service mesh (Istio/Linkerd), multi-region failover design, or contributions to open-source infrastructure tooling
Interested in this role?Continue on LinkedIn to apply.
Apply on LinkedIn