AI Infrastructure Engineer - GPU/Data Center

SproutsAI · Livermore, CA

Spotted 1h agofulltime
Job description

About this role

Employer-provided description, formatted for easier reading.

AI Infrastructure Engineer in GPU/Data Center

Livermore, CA | Hybrid (Regular Onsite)

Full-Time 4–7 Years of Experience

Only Independent Visa Holders

Strong preference for candidates with hands-on experience deploying or supporting GPU clusters in a data center environment.

We are hiring an experienced

AI Infrastructure Engineer

with strong

hands-on data center experience

to support the deployment, operation, troubleshooting, and maintenance of high-performance GPU infrastructure supporting modern AI workloads.

This is a

hands-on infrastructure role

. Candidates should have practical experience working with

physical servers, GPU systems, Linux, high-speed networking, and data center deployments

— not just cloud or software-based infrastructure.

KEY SKILLS & EXPERINCE:

  • Strong Linux Administration – RHEL, Rocky Linux, Ubuntu
  • Hands-on experience with

GPU infrastructure

– NVIDIA H100/H200/B200 or AMD Instinct

  • Experience with

InfiniBand, RoCEv2, RDMA

, and high-speed networking

  • Hands-on server hardware troubleshooting
  • Experience with

BMC, IPMI, and Redfish

  • Bare-metal provisioning, PXE, and deployment automation
  • Python and/or Bash scripting
  • Infrastructure automation and configuration
  • Data center deployment experience

, including rack installation, cabling, hardware setup, and commissioning

  • Tier 2/Tier 3 troubleshooting and root-cause analysis
  • Experience supporting production infrastructure in a data center environment

Roles and Responsibilites:

  • Deploy and maintain high-performance GPU clusters
  • Troubleshoot GPU, Linux, server hardware, and networking issues
  • Support physical data center infrastructure and production operations
  • Perform hardware installation, configuration, commissioning, and troubleshooting
  • Work with high-speed networking technologies including InfiniBand/RDMA/RoCE
  • Develop and improve infrastructure automation
  • Monitor infrastructure health, reliability, and performance
  • Participate in root-cause analysis and resolution of complex infrastructure issues

Preferred Candidates

The ideal candidate is someone who has

actually worked hands-on in a data center environment

and is comfortable working with physical servers, GPU infrastructure, Linux systems, networking, and hardware troubleshooting.

Interested in this role?Continue on SproutsAI's careers page.
Apply on SproutsAI