Software Devlopment Engineer - Kubernetes

Apple · Austin

Spotted 1h ago

What you'll need to apply

What this employer's standard application typically asks

NameEmailPhoneLocationRésuméWork authorization answer

Company-specific questions

  • Have you previously worked at Apple?
Job description

About this role

Employer-provided description, formatted for easier reading.

We're looking for a motivated Software Engineer to join our team and help build and operate reliable, secure, and scalable cloud platforms and services. In this role, you'll develop software and automation that support production systems, Kubernetes platforms, cloud infrastructure, observability, analytics, and emerging AI/ML workloads.

You'll combine software engineering skills with reliability engineering principles to improve how our platforms are built, deployed, monitored, and operated. You'll work alongside experienced engineers, architects, SREs, and AI/ML teams to solve technical challenges and continuously improve our systems.

The ideal candidate has a strong software engineering foundation, enjoys solving problems through code and automation, and is interested in cloud technologies, Kubernetes, reliability engineering, analytics, and AI.

Software Development

Design, develop, test, and maintain software, services, APIs, tools, and automation using languages such as Python and Go.

Cloud & Platform Engineering

Build and improve cloud-native services and platform capabilities across production and non-production environments.

Kubernetes

Deploy, operate, and troubleshoot applications and services running on Kubernetes. Develop automation that simplifies deployment and platform operations.

Reliability

Apply reliability engineering practices to improve system availability, performance, scalability, and operational efficiency.

Automation

Identify repetitive operational activities and develop software and automation to reduce manual effort and operational toil.

Observability

Use logs, metrics, traces, dashboards, and alerts to understand system behavior, troubleshoot issues, and identify opportunities for improvement.

Analytics

Analyze operational and application data to identify trends, anomalies, recurring issues, and performance bottlenecks. Develop dashboards and reporting that provide actionable insights.

AI/ML Infrastructure

Support infrastructure and platform capabilities for AI/ML training, LLM inference, and GPU-based workloads. Develop automation to simplify deployment and operation of these environments.

AI-Assisted Engineering

Explore and apply LLMs and AI technologies to improve software development, troubleshooting, analytics, automation, and operational workflows.

Scale & Resilience

Participate in capacity planning, performance testing, scale testing, and disaster recovery exercises.

Continuous Improvement

Identify opportunities to improve platform reliability, developer experience, automation, and operational processes.

Documentation

Create and maintain technical documentation, operational procedures, troubleshooting guides, and runbooks.

Collaboration

Work closely with software engineering, platform, SRE, QA, AI/ML, security, architecture, and program management teams.

Software Engineering

Solid understanding of software engineering fundamentals and experience developing and maintaining production software, services, tools, or automation.

Programming

Proficiency in Python and/or Go (Golang), with the ability to write clean, maintainable, and testable code.

Kubernetes

Hands-on experience with Kubernetes and containers, including deploying and troubleshooting applications. Familiarity with Helm, Kustomize, or similar tools is preferred.

Cloud

Experience with at least one major cloud platform such as AWS, Google Cloud, or Azure.

Infrastructure as Code

Familiarity with Terraform, Ansible, or similar infrastructure automation technologies.

Reliability Engineering

Understanding of reliability concepts such as monitoring, alerting, SLIs/SLOs, incident management, capacity planning, and automation.

Observability

Experience with or exposure to technologies such as Prometheus, Grafana, Splunk, OpenTelemetry, or similar observability platforms.

Analytics

Ability to analyze system and application data to identify trends and troubleshoot issues. Familiarity with SQL, Python-based data analysis, dashboards, or reporting tools is a plus.

Distributed Systems

Working knowledge of distributed system concepts including availability, scalability, networking, fault tolerance, and performance.

AI/ML

Familiarity with AI/ML concepts, LLMs, model inference, or GPU workloads is a plus. Prior AI/ML infrastructure experience is beneficial but not required.

Problem Solving

Strong analytical and troubleshooting skills with an interest in solving problems through software and automation.

Collaboration

Strong communication skills and the ability to work effectively within cross-functional engineering teams.

Interested in this role?Continue on Apple's careers page.
Apply on Apple