Benchmarking Project Lead, Siri Evaluation
What you'll need to apply
What this employer's standard application typically asks
Company-specific questions
- Have you previously worked at Apple?
About this role
Employer-provided description, formatted for easier reading.
Join the team redefining what a deeply personal and integrated assistant can be. As part of the Siri organization, you will help shape one of the world's most widely used AI assistants, powered by our next-generation of Apple Intelligence, with capabilities like personal context understanding and on-screen awareness, built with privacy from the ground up.
Your work will have direct, meaningful impact for users across iOS, iPadOS, macOS, watchOS, and visionOS. This is a rare opportunity to build at the intersection of cutting-edge AI and human-centered design, shipping technology that is centered around users and their needs.
Evaluation is at the heart of how we build our product. As Siri AI becomes more and more powerful and offers ever richer experiences to our users, our evaluations have to keep pace. We are seeking a senior manager to help drive our evaluation efforts.
The role will involve managing teams working on evaluation development and data science, and leading high-impact initiatives to bring state-of-the-art agentic evaluation to the whole Siri team.
As part of the work on next generation Siri, we are developing novel measurements of its quality. To ensure that the evaluation systems we are building are reliable, we plan benchmarking their accuracy on a wide range of features, locales, and platforms using humans in the loop.
Agentic Coding proficiency to achieve data-science, data collection and visualisation tasks Good understanding of metrics, crowd science, data collection, annotation analysis, statistics Ability to work independently and cross-functionally to integrate in partner team reporting systems and pipelines Excellent communication skills and the ability to thrive in a highly collaborative work environment
Attunement to computational linguistics, language quality, human in the loop evaluation Good engineering practices to create sustainable and easy to use data management pipelines Python experience and other tools for data collection and visualisation