Staff Experience Designer, AI Evaluation Platform
What you'll need to apply
What this employer's standard application typically asks
Company-specific questions
- Have you previously worked at Apple?
About this role
Employer-provided description, formatted for easier reading.
AI systems are only as trustworthy as the methods used to evaluate them. At Apple, where AI powers experiences for billions of people, getting evaluation right is not a support function, it is a foundational science.
Our team, part of Apple Services Engineering, builds the platform that teams across Apple use to evaluate the AI and agentic systems they ship. It's where they define what "good" means, prove it, and act on what they find. Evaluating non-deterministic systems is one of the hardest unsolved problems in production ML, and one Apple has to get right at scale.
We're looking for a Staff Experience Designer to own that experience end to end. You'll be the first designer on this platform, and evaluation as a design practice doesn't have settled patterns yet, so your work will help define them.
Teams building AI features need to know if what they shipped works. Today that means navigating unfamiliar territory: scoring something non-deterministic, telling a real regression from noise, trusting a judgment a model made instead of a person. Most existing tools here were built for the researchers who invented these methods, not for the people who now need to use them daily.
You'll set the direction here and build the design system everyone else builds on. We'd rather get a rough version in front of real users than a polished one later. And because teams use this platform to decide what to ship, trust matters more than usual: provenance, uncertainty, and edge cases determine whether someone acts on a result or quietly stops believing it.
What makes this team unusual is its interdisciplinary core. Alongside the platform, we run a research group working on evaluation methodology itself, including how to tell whether an evaluator is calibrated, biased, or measuring what it claims to. Their methods ship into this platform, and you decide how anyone first encounters them.
You'll be close to that work while it's still forming, instead of picking it up once it's finished.
8+ years of experience designing digital products or experiences, including end-to-end ownership of complex, data-dense products for technical users. You've set direction in spaces with no existing precedent, and you get to clarity by running research yourself with the people who use what you build.
A portfolio you can share, including at least one case study that walks through your process from problem framing to shipped outcome. Proven ability to design complex information: dashboards, comparison views, large result sets, and results that carry statistical uncertainty. Strong interaction and visual design skills, with a high bar for craft in dense, information-heavy interfaces.
Enough fluency in AI/ML concepts (benchmarks, metrics, model-based judging, agentic systems) to work directly with a research scientist or engineer. Fluency with AI tools in your own practice. You use Claude Code or equivalents to build working prototypes and extend what you can make on your own.
Excellent communication skills, with the ability to build buy-in across engineering and research without formal authority.
Experience as the first or only designer on a platform, and/or experience translating research output into shipped product. Experience building and maintaining a design system, including its components, patterns, and naming conventions. An AI-first instinct for interface design: shipped interfaces organized around stated intent, or surfaces meant to be operated by agents as well as people.
Experience designing coherent workflows across multiple surfaces (UI, CLI, SDK) that need to stay consistent with each other. Familiarity with modern evaluation, observability, or visualization tooling (e. g.
LangSmith, Braintrust, D3, Vega-Lite). Comfort with SQL and notebooks to explore your own data.