AIML - Manager, Applied AI Science - GenAI Model Autograding, Evaluation
What you'll need to apply
What this employer's standard application typically asks
Company-specific questions
- Have you previously worked at Apple?
About this role
Employer-provided description, formatted for easier reading.
Do you get excited by building and leading a team that redefines how AI products are evaluated? Our Evaluation organization is responsible for providing principled assessments across a diverse range of Apple features, from Search and Siri to the latest Apple Intelligence capabilities.
Within this critical function, our team specializes in leveraging advanced AI/ML techniques to enhance both the quality and efficiency of these comprehensive evaluations.
We are seeking an experienced and hands-on Manager to lead a team of Applied AI Scientists developing cutting-edge AI/ML models for the automatic grading and quality assessment of our internal GenAI products. You will set the technical direction for autograding at Apple, grow the scientists who build it, and partner across the company to make trustworthy automated evaluation a foundation that every AI product team can rely on.
In this pivotal role, you will build and lead the team that designs state-of-the-art autograder systems evaluating AI product quality at scale.
- Lead and grow a high-performing team of applied AI scientists and machine learning engineers, including hiring, talent development, performance management, and building a strong technical leadership layer.
- Set the technical vision and roadmap for autograder development, evaluation methodology, and toolings.
- Deliver reliable, scalable LLM/VLM-based evaluation systems that can be trusted for product-quality decisions.
- Navigate complex technical and organizational challenges, resolve cross-functional dependencies, and sustain alignment across organizations.
- Identify opportunities to standardize and automate the end-to-end evaluation workflow, enabling the team to scale beyond bespoke autograder development.
MS/PhD degree in Computer Science, Machine Learning, AI, or a related field.
1+ years of industry experience building LLM/VLM-based evaluators, autograders, or automated evaluation systems.
5+ years of hands-on experience in machine learning, model alignment, or benchmark creation.
Track record of managing teams of 3+ machine learning engineers and/or AI scientists.
Deep understanding of GenAI models and their evaluation methodologies, including rubric design, evaluation set development, validation, calibration, and human-model alignment.
- Strong technical judgment and a track record of leading complex, ambiguous AI/ML projects from definition through delivery.
Experience translating emerging GenAI techniques and research into reliable, production-ready systems.
Experience scaling AI/ML solutions into reusable platforms, frameworks, or standardized workflows.
Demonstrated track record of developing senior machine learning engineers and/or technical leads.
Demonstrated ability to align diverse stakeholders around strategy, navigate complex situations, and drive effective outcomes across organizations.