AI Systems Engineer - AI Platforms - Manager
Job details
- Pay
- $125,500 – $230,200 a year
- Work mode
- Hybrid
- Level
- Senior
- Experience
- 8+ years
- Education
- Bachelor's degree
- Posted
- Oct 9, 2026
- Last confirmed open
- Oct 10, 2026
About this role
Location
Anywhere in Country At EY, we’re all in to shape your future with confidence. We’ll help you succeed in a globally connected powerhouse of diverse teams and take your career wherever you want it to go. Join EY and help to build a better working world.
The opportunity
We are seeking AI Systems Engineers to build and operate the foundational substrate that powers EY’s AI-native platform. This role owns the infrastructure and cloud-native platform layers of the Hybrid AI Multi-Environment Runtime (HAI), from bare-metal and GPU infrastructure through Kubernetes, cluster fabric, and multi-tenant scaling.
You will be responsible for building and managing EY Fabric environments across cloud, on-prem, edge, and air-gapped targets. This is the substrate on which EY Agentic AI capabilities run.
This role is ideal for a full-stack infrastructure leader who is equally comfortable with bare-metal and GPU systems and production Kubernetes at scale, who treats reliability and portability as non-negotiable in regulated client contexts, and who understands that the substrate is a product in its own right, measured by the velocity, safety, and portability it unlocks for every team building above it.
Your Key Responsibilities
Own the cluster & cloud-native platform: compute, Kubernetes and scheduling, cluster fabric/networking, multi-tenancy, and distributed compute, as the substrate for Agentic AI workflows and tooling.
Own the infrastructure foundation
Ubuntu/OS, BMC/bare-metal, DPU architecture, and NVAIE (GPU/Network/DCGM), ensuring the physical and virtual bedrock is provisioned, patched, and production-ready. Stand up and manage EY Agentic AI environments across cloud (EKS/AKS/GKE), on-prem AI Factory (RKE2/NVAIE), edge (K3s), and air-gapped deployment modes, maintaining one consistent stack contract across all targets.
Deliver foundational platform capabilities such as Infrastructure Management, Kubernetes & Scheduling, and Cluster Fabric Management, so downstream runtime, data, and execution services can run safely and consistently. Own cluster lifecycle, autoscaling, GPU pooling/virtualization, and multi-tenancy boundaries (vCluster/Crossplane/Karpenter), providing isolated, elastic capacity per tenant and engagement.
Own secure execution and inference
Ray Serve, vLLM/NIM/Triton, and NVIDIA Dynamo, with sandboxed execution (gVisor/Firecracker for hosted, NVIDIA OpenShell/vNode for on-prem) for isolated, safe model execution.
Own cognitive and routing
Envoy AI Gateway, semantic routing (vLLM-SR), model/prompt selection, and streaming response handling — directing each request to the right model under the right constraints. Collaborate with DevOps Engineers on deployment and delivery of the platform itself: CI/CD/CV (ArgoCD), infrastructure-as-code / GitOps (Helm/OpenTofu), so environments are reproducible and drift-free.
Own backup, disaster recovery, and cross-environment replication for high availability (Velero, CloudNativePG, Cilium ClusterMesh), along with patching and platform supply-chain hygiene. Ensure the substrate is modular and swappable, so components can be replaced without rewriting consumers, minimizing vendor lock-in while preserving the stack contract.
Skills And Attributes For Success Deep expertise in cloud-native platform engineering (Kubernetes, networking, multi-tenancy at scale) and low-level infrastructure and systems engineering (bare-metal, GPU, DPU). Advanced understanding of how compute, networking, storage, and scheduling interact to form a reliable, portable substrate.
Comfortable operating across cloud, on-prem, edge, and air-gapped environments simultaneously, with a portability-first mindset. Strong command of AI gateways, semantic routing, and model/prompt selection under latency and cost constraints. A passion for ensuring reliability, repeatability, and reduction of operational toil through automation and infrastructure-as-code.
Ability to define and honor clean ownership boundaries with adjacent trust, data, and runtime teams. Strong communicator able to explain infrastructure tradeoffs to engineers, architects, and leadership. Natural product-ownership orientation toward foundational platforms, measured by the downstream velocity and reliability they enable.
To qualify you must have Bachelor’s or Master’s degree in Computer Science or related technical field. 8+ years building or operating enterprise infrastructure, cloud platforms, or large-scale Kubernetes environments, including hands-on systems depth. Hands-on expertise with Kubernetes distributions (RKE2, EKS/AKS/GKE, K3s) and full cluster lifecycle management.
Deep experience with bare-metal, cloud, hybrid, on-prem, and ideally air-gapped deployment models. Strong grounding in cluster networking (Cilium/service mesh/CNI), storage, and multi-tenancy isolation. Experience with GPU infrastructure and scheduling (NVAIE/DCGM, GPU operators, virtualization/MIG).
Experience with infrastructure-as-code and GitOps tooling (Terraform/OpenTofu, Helm, ArgoCD, Crossplane). Proven track record operating production infrastructure under compliance, security, or regulatory constraints. Ability to work effectively with security, architecture, product, and delivery teams.
Ideally, you’ll also have Experience supporting AI/ML and GPU-intensive workloads (GPU virtualization, MIG, AI workload scheduling). Experience with multi-tenant platforms and tenancy isolation at scale (e.g., vCluster). Experience with disaster recovery, backup, and cross-environment replication tiered by RPO/RTO.
Familiarity with distributed compute and batch orchestration (Kueue, Slurm/Slinky, Ray, OpenMPI). Background in platform or SRE roles where success is measured by downstream team velocity, reliability, and portability. Exposure to regulated delivery environments (financial services, tax, healthcare, risk).
What We Offer
- You At EY, we’ll develop you with future-focused skills and equip you with world-class experiences.
- We’ll empower you in a flexible environment, and fuel you and your extraordinary talents in a diverse and inclusive culture of globally connected teams.
- Learn more.
- We offer a comprehensive compensation and benefits package where you’ll be rewarded based on your performance and recognized for the value you bring to the business.
- The base salary range for this job in all geographic locations in the US is $125,500 to $230,200.
- The base salary range for New York City Metro Area, Washington State and California (excluding Sacramento) is $150,700 to $261,600.
- Individual salaries within those ranges are determined through a wide variety of factors including but not limited to education, experience, knowledge, skills and geography.
- In addition, our Total Rewards package includes medical and dental coverage, pension and 401(k) plans, and a wide range of paid time off options.
- Join us in our team-led and leader-enabled hybrid model.
- Our expectation is for most people in external, client serving roles to work together in person 40-60% of the time over the course of an engagement, project or year.
- Under our flexible vacation policy, you’ll decide how much vacation time you need based on your own personal circumstances.
- You’ll also be granted time off for designated EY Paid Holidays, Winter/Summer breaks, Personal/Family Care, and other leaves of absence when needed to support your physical, financial, and emotional well-being.
- Are you ready to shape your future with confidence?
- Apply today.
- EY accepts applications for this position on an on-going basis.
- For those living in California, please click here for additional information.
- EY focuses on high-ethical standards and integrity among its employees and expects all candidates to demonstrate these qualities.
- EY | Building a better working world EY is building a better working world by creating new value for clients, people, society and the planet, while building trust in capital markets.
- Enabled by data, AI and advanced technology, EY teams help clients shape the future with confidence and develop answers for the most pressing issues of today and tomorrow.
- EY teams work across a full spectrum of services in assurance, consulting, tax, strategy and transactions.
- Fueled by sector insights, a globally connected, multi-disciplinary network and diverse ecosystem partners, EY teams can provide services in more than 150 countries and territories.
- EY provides equal employment opportunities to applicants and employees without regard to race, color, religion, age, sex, sexual orientation, gender identity/expression, pregnancy, genetic information, national origin, protected veteran status, disability status, or any other legally protected basis, including arrest and conviction records, in accordance with applicable law.
- EY is committed to providing reasonable accommodation to qualified individuals with disabilities including veterans with disabilities.
- If you have a disability and either need assistance applying online or need to request an accommodation during any part of the application process, please call 1-800-EY-HELP3, select Option 2 for candidate related inquiries, then select Option 1 for candidate queries and finally select Option 2 for candidates with an inquiry which will route you to EY’s Talent Shared Services Team (TSS) or email the TSS at [email protected].