Application Reliability Engineer
What you'll need to apply
What this employer's standard application typically asks
Company-specific questions
- Have you previously worked at Apple?
About this role
Employer-provided description, formatted for easier reading.
At Apple, great ideas have a way of becoming great products, services, and customer experiences very quickly. Bring passion and dedication to your job and there's no telling what you could accomplish.
Payments and Financial Engineering builds the systems that power Apple's Finance and Accounting capabilities, processing very high-volume micro-transactions across Apple Pay, App Store, iTunes, Retail, Online, and Reseller channels to ensure accurate invoicing, disbursements, receipts, and payments.
We are looking for an Application Reliability Engineer (ARE) to own the operational health of this platform - driving problem management, root cause resolution, and automation across a distributed, high-scale system that underpins a majority of Apple's revenue.
In this role, you will be a strong software engineer with production-debugging and distributed-systems expertise, responsible for the reliability of a mission-critical payments platform - not simply triaging and closing incidents, but diagnosing, automating, and permanently eliminating the underlying application issues.
You will reproduce failures, read and modify production code to implement the fix, and add detection or prevention mechanisms so the same class of issue doesn't recur.
5-7 years of experience in software engineering, application support, or SRE roles for critical, high-scale production systems, ideally in a payments, financial, or transaction-processing domain. Strong programming skills in Java or another JVM language, with working proficiency in at least one of Python, Go, or C++, or Shell Scriptng.
Ability to read, debug, and modify production code to implement permanent fixes, not just workarounds. SQL and database troubleshooting experience (e. g.
, Oracle, MongoDB, PostgreSQL), including diagnosing connection pool exhaustion, stale connections, and slow queries.
Practical understanding of concurrency, threading, and asynchronous processing, and the failure modes they introduce (thread/consumer starvation, blocked threads, race conditions) Experience with observability stacks - logs, metrics, and traces - to diagnose application-level observability gaps and reproduce failures Strong Linux/Unix fundamentals, including process management, file permissions, and cron Performance profiling and troubleshooting experience, including identifying resource leaks and dependency timeouts Experienced in troubleshooting and driving root cause resolution of production incidents, including identifying performance bottlenecks and proposing fixes Self-starter with strong customer and product focus, and a desire to gain awareness of an ecosystem and how frontend and backend systems collaborate.
Ability to communicate thoughtfully, write engineering proposals and support playbooks, leverage problem-solving skills, build a learning mindset and establish long-term relationships.
Experience with containerization and orchestration (Docker, Kubernetes) Familiarity with CI/CD systems and infrastructure fundamentals across prod/non-prod environments Experience with cloud platforms (AWS, Azure, GCP), including managed services and cloud migration Experience with caching systems (Redis, Memcached) and messaging/queueing systems (Kafka, SQS, RabbitMQ), including queue offset management and retry/dead-letter handling.
Working knowledge of system design principles applicable to high-throughput, high-availability platforms. Experience with distributed storage systems (S3, GridFS). Track record of building tools or frameworks that detect and prevent recurring classes of production issues.
Experience setting up trend analysis or observability processes to proactively identify chronic issues. Solid grounding in REST/API and distributed-systems fundamentals Contributions to system design reviews or architecture discussions focused on reliability and scalability.