Site Reliability Engineer

Evlo AI · Miami, FL

Spotted 2h agofulltime
Job description

About this role

Employer-provided description, formatted for easier reading.

About The Role

The role keeps production systems running at scale — owning the reliability, observability, and infrastructure that powers services handling millions of requests per day.

You will work closely with software engineers and platform teams to eliminate toil, automate operations, and design systems that fail gracefully instead of catastrophically.

Key Responsibilities

  • Build and maintain Kubernetes-based infrastructure across multiple cloud environments (AWS or GCP), including clusters, ingress, autoscaling policies, and workload orchestration
  • Design and implement observability stacks using Prometheus, Grafana, and OpenTelemetry — dashboards, SLOs, and alerting that catch real incidents without generating noise
  • Lead incident response for production outages: triage, mitigation, root cause analysis, and blameless postmortems with concrete follow-up actions
  • Automate infrastructure provisioning and configuration with Terraform, Ansible, and CI/CD pipelines (GitHub Actions or GitLab CI) to make deployments boring and repeatable
  • Reduce toil through automation — writing tooling in Python or Go that eliminates manual operational work
  • Partner with development teams on capacity planning, performance tuning, and reliability reviews for new services before they hit production
  • Contribute to chaos engineering practices, failure testing, and disaster recovery runbooks for critical systems

What We Are Looking For

  • 3–7 years of experience in SRE, DevOps, or production infrastructure engineering
  • Deep hands-on experience with Kubernetes in production — deployment, debugging, RBAC, networking, and scaling
  • Strong proficiency in at least one scripting or systems language: Python, Go, or Bash
  • Production experience with infrastructure-as-code (Terraform) and modern CI/CD pipelines
  • Solid grasp of distributed systems fundamentals: load balancing, caching, queuing, replication, and failure modes
  • Experience running on-call and leading incident response for customer-facing services
  • Bachelor's degree in Computer Science, Engineering, or equivalent practical experience
  • Bonus: Experience with service meshes (Istio, Linkerd), chaos engineering tools (Gremlin, Litmus), cloud security posture tooling, or contributing to open-source infrastructure projects
Interested in this role?Continue on Evlo AI's careers page.
Apply on Evlo AI