AIML - Sr Software Engineer - AI, Evaluation

Apple · Cupertino

Spotted 1h ago

What you'll need to apply

What this employer's standard application typically asks

NameEmailPhoneLocationRésuméWork authorization answer

Company-specific questions

  • Have you previously worked at Apple?
Job description

About this role

Employer-provided description, formatted for easier reading.

Do you want to help build the next generation of Apple AI products? Our Evaluation organization measures quality across Apple Intelligence, Siri, and newest generative features, and our measurements directly inform launch decisions.

Our team builds the platforms behind it

LLM-as-judge autograders, tools to validate them against human judgment, and systems that run them at scale. We are looking for a senior engineer to own a critical part of that platform and take it to the next level.

This is an ownership role. You will take on major systems in our evaluation platform, existing or new, and be accountable for its direction, quality, and adoption. This role sits at the intersection of AI modeling, software engineering, and product quality.

We are expanding rapidly, so the ability to build a strong foundation and structure to any system is paramount. We are supporting features that encompass both product development and research, which requires a balance of respect for both the scientific process as well as the product release lifecycle. Our systems connect many platforms and serve teams across Apple, so judgment with people matters as much as technical depth.

BS/MS/PhD in Computer Science, Machine Learning, or a related field. 8+ years of software engineering experience, including owning production systems used by other teams. Exceptional Python skills and a strong engineering understanding of system design, API design, system testing and validation, debugging and monitoring.

Experienced in shipping production code with AI coding agents while keeping quality high.

Proficient in building agentic systems, LLM-based applications, LLM-as-judge systems, and offline evaluations. Ability to interpret evaluation results, including human agreement, run-to-run consistency, and respecting constraints like data retention and privacy policies.

Expert in building and maintaining technical integrations across multiple platforms in a large organization, with varying technical requirements and needs based off scaling, volume, throughput, etc. Product-minded, turning ambiguous requirements into a focused plan.

Interested in this role?Continue on Apple's careers page.
Apply on Apple