Software Engineer, Triage Services
What you'll need to apply
What this employer's standard application typically asks
Company-specific questions
- Have you previously worked at Apple?
About this role
Employer-provided description, formatted for easier reading.
The System Triage Services team builds mission-critical applications that help detect, analyze, and classify low-level software crashes across all Apple platforms. You will design and build the next generation of AI-powered triage solutions for Core OS - work that directly supports the reliability of over 2 billion active Apple devices.
This role blends distributed systems engineering with hands-on agent development, tackling problems in reliability, scale, and system design alongside applied AI. If you want to understand operating systems at a deep level and are motivated by shipping work with measurable, visible impact, this role is for you!
You will design and build agentic systems that transform how Core OS triages software bugs and features at scale — building production-grade infrastructure that runs AI/ML workloads reliably and improves over time through rigorous evaluation and feedback loops. The goal is to make autonomous triage a dependable, first-class part of how Core OS ships and debugs software.
You'll work across the build and integration lifecycle, agent orchestration, and the infrastructure that connects them, partnering with project managers and engineering teams across Software, Hardware, and Silicon groups to turn complex triage requirements into resilient systems.
Programming experience in Python. 5+ years of experience building scalable data platforms and distributed systems. Experience with LLM and agent development.
Knowledge of RAG architectures, embedding generation, vector databases, and AI data preparation for agentic workflow. Strong collaboration and communication skills working across multi-domain software engineering teams. Familiarity with eval frameworks for agent quality (offline benchmarks, drift detection, confidence scoring).
Bachelor's degree in Computer Science, Data Engineering, or a related field.
Master's degree in Computer Science, Data Engineering, or a related field. 1+ year developing production-grade agentic systems. Experience with large-scale telemetry or observability systems feeding automated decision-making.
Experience shipping tools adopted by engineering teams beyond your own. Background in kernel or OS-level debugging, crash analysis, or root-cause investigation.