Back to jobs

Site Reliability Engineer

Evlo AI · New York, NY

Spotted 13m agoFull-time

Job details

Employment
Full-time
Level
Mid level
Experience
3+ years
Education
Bachelor's degree
Posted
Oct 11, 2026
Last confirmed open
Oct 11, 2026
Job description

About this role

About The Role

The role owns the reliability of large-scale production systems — distributed services, Kubernetes-based infrastructure, and the pipelines that deploy, observe, and heal them. SLOs, error budgets, and incident response are not buzzwords here; they are the operating model.

You will join a platform reliability team that partners directly with product engineering groups, ensuring that latency, availability, and infrastructure cost targets hold under real traffic — millions of requests per day — while the platform continues to ship.

Key Responsibilities

  • Design, implement, and maintain SLOs and error budgets for critical services, driving prioritization of reliability work across engineering teams
  • Build and evolve observability tooling — Prometheus, Grafana, OpenTelemetry, distributed tracing — so on-call engineers can diagnose issues in minutes, not hours
  • Automate infrastructure provisioning and configuration with Terraform, Ansible, and GitOps workflows on Kubernetes (EKS/GKE)
  • Lead and participate in incident response: drive mitigations, write blameless postmortems, and track remediation items to completion
  • Improve deployment pipelines — CI/CD with GitHub Actions or GitLab CI, canary and blue-green strategies, automated rollback — to reduce deployment risk without slowing release cadence
  • Perform capacity planning, load testing, and performance tuning for stateful and stateless services across multi-region cloud environments (AWS or GCP)
  • Author runbooks and automation that eliminate toil, progressively shifting pager load from humans to self-healing systems

What We Are Looking For

  • 3–6 years of experience in SRE, DevOps, or backend infrastructure engineering, including running production systems at scale with on-call responsibility
  • Strong proficiency with Kubernetes in production: cluster operations, Helm, autoscaling, networking, and troubleshooting under pressure
  • Deep hands-on experience with at least one major cloud provider (AWS, GCP, or Azure), including IAM, VPC networking, and managed data services
  • Solid coding skills in Python, Go, or Bash for building tooling, automation, and infrastructure-as-code (Terraform required)
  • Proven experience with monitoring and observability stacks (Prometheus/Grafana, Datadog, or equivalent) and defining meaningful service-level metrics
  • Track record of leading incident response and writing clear, actionable postmortems; strong written communication is essential
  • Bachelor's degree in Computer Science, Engineering, or equivalent practical experience. Bonus: experience with chaos engineering, service mesh (Istio/Linkerd), multi-region failover design, or contributions to open-source infrastructure tooling
Interested in this role?Continue on LinkedIn to apply.
Apply on LinkedIn