Software Engineer Manager: Agentic Evaluation Platform

Apple · Cupertino

Spotted 2h ago

What you'll need to apply

What this employer's standard application typically asks

NameEmailPhoneLocationRésuméWork authorization answer

Company-specific questions

  • Have you previously worked at Apple?
Job description

About this role

Employer-provided description, formatted for easier reading.

Join the team redefining what a deeply personal and integrated assistant can be and lead the platform that defines how we measure it.

As part of the Siri Agentic Evaluation Engineering organization, you will help shape one of the world's most widely used AI assistants, powered by our next generation of Apple Intelligence, with capabilities like personal context understanding and on-screen awareness, built with privacy from the ground up. Your work will have direct, meaningful impact for users across iOS, iPadOS, macOS, watchOS, and visionOS.

As the engineering leader of the evaluation platform and services, you'll own the systems that run and scales Siri’s evaluations, leveraging the best Apple platform capabilities and unlocking new ones to create novel evaluations. Your team applies large scale infrastructure, apple platforms and service, and AI agents to steer, validate, and judge evaluation scenarios at scale.

These generate the signals that leadership relies on to make ship decisions. When we say a capability is ready, it's your platform that backs the claim.

This is a rare opportunity to build scalable software and services for real world products at the intersection of cutting-edge AI and human-centered design, shipping technology that is centered around users and their needs.

In this role you'll lead and grow a team of engineers building the backend services, compute orchestration, and evaluation harnesses that let us evaluate Siri reliably and at scale. You'll own the technical direction and the roadmap, and you'll be accountable for delivery.

The through-line of the work is interfaces. Your team owns the APIs that expose evaluation as a platform capability, the harness contracts that determine how devices and models are driven through a run, and the data workflows that carry results from execution through scoring to reporting.

Those boundaries are the product: they decide how much compute the organization can actually use, how quickly a failure can be traced to its cause, and how many teams can build on evaluation without coming to you first. Get them right and the platform scales past your team; get them wrong and everything routes through you.

You'll also have meaningful autonomy in how you get there. Both the evaluation platform and the underlying Siri architecture are evolving, so the specific frameworks, runtimes, and components will change over time.

We're looking for a leader who operates well in that ambiguity, someone who can reshape team scope deliberately as the domain matures, decide what to invest in versus retire, and keep the team's momentum through the change.

Just as importantly, you'll build the team itself: recruiting, mentoring, and developing engineers including senior ICs, with a focus on engineering rigor, technical depth, and clear ownership.

5+ years of experience developing production software (e. g.

, Python, Java, Go, or Swift) 3+ years of engineering management experience, including owning roadmap and delivery, mentoring, setting clear expectations, and providing consistent, honest feedback Strong technical leadership and the ability to set architectural direction and inspire a team Experience with test infrastructure, CI/CD, or evaluation/ML platforms Proactive and self-motivated, with demonstrated creative and critical-thinking abilities Excellent spoken and written communication skills, and the ability to collaborate across teams M.

S. or B. S.

in Computer Science, Machine Learning, or a related field, or equivalent experience

Experience building or leading teams that apply LLMs / AI agents to real engineering workflows (e. g.

, data generation, code generation, or agentic pipelines with human-in-the-loop review) Depth in one or more of: distributed systems and backend services (REST/gRPC), cloud infrastructure (Docker/Kubernetes), data platforms, or large-scale test and release infrastructure Familiarity with LLM and Agentic evaluation concepts Experience implementing model evaluation frameworks and experiment workflows, including offline and online experiments, to measure how model or platform changes affect key metrics such as accuracy, latency, and engagement Experience building machine learning platforms and infrastructure used by multiple teams for experimentation and deployment Experience designing and integrating monitoring, logging, and alerting solutions — dashboards, anomaly detection, and real-time alerts — to track reliability and performance in production systems Experience developing data processing pipelines with big data frameworks (such as Apache Spark, Kafka, or Flink) to ingest, transform, and analyze large-scale log or event data Demonstrated cross-functional leadership and the ability to drive coverage and quality decisions with senior stakeholders Experience hiring and growing a team, including senior engineers Comfort operating in ambiguity and reshaping team scope as the domain matures

Interested in this role?Continue on Apple's careers page.
Apply on Apple