Site Reliability Engineer (SRE)
Job details
- Pay
- $145,000 – $225,000 a year
- Level
- Entry level
- Experience
- 2+ years
- Posted
- Oct 7, 2026
- Last confirmed open
- Oct 8, 2026
About this role
Location
New York City, NY
Experience: 2–8 years
Employment Type: Full-time
Compensation: $145,000 – $225,000 per year + bonus/equity
About the Role
We are looking for a highly skilled Site Reliability Engineer (SRE) to join our engineering team in New York City. You will be responsible for ensuring the reliability, scalability, performance, and security of our production systems and cloud infrastructure.
You will work at the intersection of software engineering, cloud infrastructure, DevOps, and operations. The ideal candidate is passionate about automation, observability, distributed systems, and building highly reliable platforms that can scale with business growth.
Requirements
Key Responsibilities
- Design, build, and maintain highly available and scalable production infrastructure.
- Improve system reliability, performance, scalability, and operational efficiency.
- Develop automation to eliminate repetitive operational tasks and reduce manual intervention.
- Manage and optimize cloud infrastructure across
AWS, GCP, or Azure
.
- Build and maintain CI/CD pipelines for reliable and automated deployments.
- Manage containerized workloads using
Docker and Kubernetes
.
- Implement infrastructure using
Terraform, CloudFormation, or similar Infrastructure-as-Code tools
.
- Establish monitoring, logging, alerting, and observability across production systems.
- Work with tools such as
Prometheus, Grafana, Datadog, CloudWatch, ELK, or OpenTelemetry
.
- Troubleshoot production incidents and participate in incident response and root-cause analysis.
- Define and monitor
SLIs, SLOs, and SLAs
.
- Identify system bottlenecks and implement performance and reliability improvements.
- Develop disaster recovery, backup, and business continuity strategies.
- Collaborate with software engineers to build reliable and operationally efficient applications.
- Participate in on-call rotations and help improve incident-management processes.
- Document infrastructure, operational procedures, architecture, and incident learnings.
Must-Have Skills
- 2–8 years of experience
in SRE, DevOps, Cloud Engineering, Infrastructure Engineering, or a related field.
- Strong experience with
Linux/Unix systems administration
.
- Hands-on experience with
AWS, GCP, or Azure
.
- Strong knowledge of
Docker and Kubernetes
.
- Experience with
Terraform or other Infrastructure-as-Code tools
.
- Proficiency in at least one programming or scripting language such as
Python, Go, Bash, or Java
.
- Experience building and managing
CI/CD pipelines
.
- Strong understanding of networking, DNS, TCP/IP, load balancing, and security fundamentals.
- Experience with monitoring, logging, alerting, and observability platforms.
- Understanding of distributed systems, scalability, high availability, and fault tolerance.
- Strong troubleshooting and problem-solving skills.
- Experience with incident response and root-cause analysis.
Nice-to-Have Skills
- Experience with
AWS EKS, ECS, EC2, Lambda, RDS, S3, CloudFront, and CloudWatch
.
- Experience with
Prometheus, Grafana, Datadog, New Relic, or OpenTelemetry
.
- Knowledge of service meshes such as Istio or Linkerd.
- Experience with
Kafka, RabbitMQ, or other distributed messaging systems
.
- Experience with GitHub Actions, GitLab CI, Jenkins, or ArgoCD.
- Knowledge of security engineering and cloud security best practices.
- Experience implementing SLOs, error budgets, and reliability metrics.
- Familiarity with FinOps and cloud-cost optimization.
- Experience operating large-scale distributed systems.
- AWS/GCP/Azure or Kubernetes certifications.
- Experience in fintech, SaaS, high-growth startups, or other high-availability environments.