Platform Engineer - DevOps Specialist L2
About this role
Employer-provided description, formatted for easier reading.
Remote
Contract (7 months 17 days)
Published 5 hours ago
aws
EKS
trainium
kubernetes
inferentia
LLM Platforms
Role Summary:
We are seeking an experienced AI Infrastructure Engineer to support and optimize large-scale AI/ML platforms within AWS environments. This role focuses on AI model training and inference infrastructure, distributed training workloads, cloud-native platform engineering, and AWS AI accelerator technologies including Trainium and Inferentia.
The ideal candidate will have experience building scalable AI infrastructure and supporting modern machine learning workloads at scale.
Key Responsibilities:
- Design, deploy, and support AI/ML infrastructure within AWS environments.
- Build and maintain scalable platforms for large-scale model training and inference workloads.
- Support distributed training environments for LLMs and Generative AI workloads.
- Implement and manage cloud-native infrastructure using Kubernetes, EKS, and containerized technologies.
- Utilize AWS Trainium, Inferentia, EC2 Trn instances, and AWS Neuron SDK to support AI workloads.
- Optimize AI platform performance through benchmarking, tuning, and infrastructure improvements.
- Support migration and optimization efforts involving AWS AI accelerator technologies.
- Collaborate with engineering teams to deliver scalable and reliable AI infrastructure solutions.
Required Qualifications:
- 5+ years of experience in Cloud, Data, or AI Infrastructure.
- Experience supporting large-scale AI/ML training workloads within AWS environments.
- Experience with AWS Trainium, Inferentia, EC2 Trn instances, and AWS Neuron SDK for AI model training and inference.
- Strong understanding of LLMs, Generative AI, distributed training, and AI/ML infrastructure.
- Hands-on experience with Kubernetes (EKS), Docker, Python, PyTorch, and cloud-native architectures.
- Experience with AI performance optimization, benchmarking, and infrastructure scalability.
Preferred Qualifications:
- Exposure to AWS Trainium or Inferentia environments.
- Experience optimizing or migrating AI workloads to AWS AI accelerator platforms.
- Experience supporting large-scale distributed AI training environments.
The pay range that the employer in good faith reasonably expects to pay for this position is $37.37/hour - $58.39/hour. Our offered benefits include medical, dental, vision and retirement benefits. Applications will be accepted on an ongoing basis.
Tundra Technical Solutions is among North America’s leading providers of Staffing and Consulting Services. Our success and our clients’ success are built on a foundation of service excellence.
We are an equal opportunity employer, and we do not discriminate on the basis of race, religion, color, national origin, sex, sexual orientation, age, veteran status, disability, genetic information, or other applicable legally protected characteristic.
Qualified applicants with arrest or conviction records will be considered for employment in accordance with applicable law, including the Los Angeles County Fair Chance Ordinance for Employers and the California Fair Chance Act.
Unincorporated LA County workers: we reasonably believe that criminal history may have a direct, adverse and negative relationship with the following job duties, potentially resulting in the withdrawal of a conditional offer of employment: client provided property, including hardware (both of which may include data) entrusted to you from theft, loss or damage; return all portable client computer hardware in your possession (including the data contained therein) upon completion of the assignment, and; maintain the confidentiality of client proprietary, confidential, or non-public information.
In addition, job duties require access to secure and protected client information technology systems and related data security obligations.