AIML - Machine Learning Research Lead, RL Agents, MLR
What you'll need to apply
What this employer's standard application typically asks
Company-specific questions
- Have you previously worked at Apple?
About this role
Employer-provided description, formatted for easier reading.
We are looking for a hands-on research lead to drive our work on reinforcement learning and post-training for agentic AI, and to manage a small team of senior researchers working on related problems in RL, agentic tool-calling, synthetic environment generation, model scaling, and multimodal action models.
You will help set direction for how we develop infrastructure, training, runtime and evaluation procedures for interactive agents — tool calling, coding, computer use, and long-horizon tasks.
This role sits inside a research organization pursuing first-principles approaches to core AI problems: generative foundation models across modalities (text, images, graphs, scientific and engineering data), vision-language modeling and implicit world modeling, self-supervised learning, and search and evolutionary methods for optimizing both agents and the environments they learn in.
A distinctive part of our agenda is designing methods that fit Apple's deployment reality — on-device and hybrid (device plus private cloud) execution, co-designed with current and future hardware — and that take advantage of what this ecosystem uniquely enables, such as deeply personalized, long-context agentic experiences.
We aim for both field-changing research and direct impact on Apple products and internal engineering processes.
MLR is a research group first. Management here is about spreading the load of running a team, not stepping away from the work — everyone, including leads, stays hands-on. We support continued engagement with the academic community: publishing, conference service, student collaboration, and internships.
- Lead research on RL and post-training for agentic capabilities: reward, preference optimization, and verifier design, training recipes, and evaluation for tool calling, coding, and multi-step interactive tasks.
- Build and own synthetic data and task-generation pipelines — generating diverse, verifiable tasks and environments, along with the interactive environments and benchmarks that go with them, and the curricula that turn them into capable agents.
- Drive codebases and infrastructure for the core RL research effort and help engage partner teams to use and co-develop the framework.
- Manage and mentor a small team (3–4) of senior researchers and research engineers with distinct specialties, shaping a shared research direction while protecting room for bottom-up, idea-driven work.
- Stay hands-on: run experiments, write code, and contribute directly to the team's most important technical problems.
- Connect post-training research to efficiency and deployment: what works under on-device and hybrid compute constraints, and how method design interacts with hardware.
- Collaborate across the organization on adjacent directions, including methods for environment and agent co-optimization, self-improvement, world models used as planners or policies, and personalized long-context agents.
- Publish in top venues and engage with the broader research community.
PhD in machine learning or a related field, or equivalent research experience
7-10+ years of research experience beyond PhD in industry or as an academic research lead
Strong track record in RL and/or post-training of large models, demonstrated through publications, open-source contributions, or shipped systems
Leadership experience: setting and defending a research direction over multiple years, and directing others' work — through direct reports, PhD students, postdocs, or sustained project teams. Formal management experience is welcome but not required
Experience owning ML infrastructure, frameworks and codebases, including open-source research frameworks or environment suites others build on
Strong engineering skills; comfortable working hands-on in large training codebases
Experience taking research from idea to product or production impact
Familiarity with efficiency-aware modeling: small models, mixture-of-experts, quantization, distillation, inference-cost constraints, or hardware-aware method design
Interest or background in open-endedness, evolutionary computation, curriculum or environment design, multi-agent systems, or self-improving systems
Principled or theoretical grounding in RL — representation, exploration, or optimization views of policy learning — alongside strong empirical work
Breadth across core machine learning — generative models, self-supervised learning, pre-training — and perspective on the field's longer arcs, not only its most recent methods
Experience growing other researchers, and managing researchers and engineers with heterogeneous specialties and synthesizing their work toward a common goal
Experience owning a large RL or post-training codebase