Site Reliability Engineer / Devops — Retail Engineering
What you'll need to apply
What this employer's standard application typically asks
Company-specific questions
- Have you previously worked at Apple?
About this role
Employer-provided description, formatted for easier reading.
Apple’s IS&T Retail Engineering team is seeking a SDET to own quality across our retail ecosystem — spanning eCommerce backend services, SAP integrations, and cross-functional end-to-end workflows. This is a hands-on technical leadership role: you’ll architect automation frameworks, lead testing strategy across multiple concurrent projects, and drive quality standards organization-wide.
You’ll work at the intersection of backend development, test automation, and program coordination — ensuring that features shipping to apple. com/shop and supporting retail systems meet Apple’s bar for reliability, performance, and customer experience.
You will drive the reliability, deployment, and scalability of compute platforms across on-premises and hybrid cloud environments. Collaborating closely with cross-functional technical and business partners, you will build Infrastructure as Code, optimize container orchestration, and streamline CI/CD delivery pipelines.
You will champion automation and operational excellence, ensuring high availability, robust security standards, and proactive observability across large-scale distributed systems.
Bachelor’s degree in Computer Science or equivalent field with 7+ years of experience, or Master’s degree with 5+ years of experience. 7+ years of experience in Site Reliability Engineering with a strong focus on building, scaling, and operating large-scale distributed platform services, and Java applications. Strong technical grasp of Open Source technologies designed for large-scale data processing.
Proven expertise in designing, analyzing, and troubleshooting complex distributed systems. Proficiency in at least one modern programming or scripting language (Python, Java, Go, Bash, Ansible, or similar). Practical experience designing and deploying end-to-end observability stacks (Prometheus, Grafana, Datadog, OpenTelemetry, ELK, etc.)
Demonstrated troubleshooting and problem-solving skills across production software and infrastructure environments.
In-depth understanding of SRE principles, including error budgeting, SLO/SLI/SLA definition, and advanced observability practices (Prometheus, Splunk, Grafana, OpenTelemetry). Advanced programming skills in Java, Python, or Go, with hands-on experience across relational, NoSQL, or OLAP databases and event-driven streaming architectures (Kafka, RabbitMQ).
Track record of managing production on-call rotations, critical incident triage, root cause analysis (RCA), and post-incident reviews (PIR). Solid knowledge of enterprise security standards, cryptography, authentication protocols (OAuth, SAML, SSO), and compliance governance.