Memory Subsytem Expert, MLA Technology

Amazon · Austin, Texas, USA

Spotted 1h ago

What you'll need to apply

What this employer's standard application typically asks

NameEmailPhoneLocationRésuméWork authorization answerVisa sponsorship answer

Company-specific questions

  • Have you worked for Amazon in the past?
Job description

About this role

Employer-provided description, formatted for easier reading.

Annapurna Labs (our organization within AWS UC) designs silicon and software that accelerates innovation. Customers choose us to create cloud solutions that solve challenges that were unimaginable a short time ago—even yesterday. Our custom chips, accelerators, and software stacks enable us to take on technical challenges that have never been seen before, and deliver results that help our customers change the world.

In Annapurna Labs we are at the forefront of hardware/software co-design not just in Amazon Web Services (AWS) but across the industry. Our Machine Learning Accelerator (MLA) Technology is seeking a memory subsystem expert who is interested in diving deep into the definition, design, validation, and data center operation of AWS’s next generation machine learning silicon and servers.

Our team spans HBM, SRAM, and chiplet interconnects to enable the highest performing memory subsytems in the world.

As a senior member of our technology team, you will have opportunities to participate in the design and execution of HBM, SRAM, UCIe, and general high speed analog technologies, with the goal of creating the most stable machine learning platforms within AWS’s data centers.

A senior memory subsytem engineer on our team needs to be able to work with vendors and internal design teams, understand complex memory and SERDES features, write/modify tests at scale, debug fleet wide issues, and collect data from manufacturing and the data center.

Our broader team has end to end ownership of some of the most complicated IPs on the most advanced server hardware in the world. We drive complex technical debug efforts involving our IPs and leverage the massive scale of EC2 to monitor, optimize, and improve our machine learning hardware reliability on behalf of our customers.

Key job responsibilities

As a senior member of the team, you will join a mixed group of hardware and software engineers working to design, integrate, and innovate the next generation of machine learning chips into Trainium servers. In this position it is expected that you will:

  • Collaborate with architects, design teams, and software engineers on our next generation ML chips
  • Support on-going debug and operations of previous ML chips within manufacturing and the data center
  • Dive deep into IP integration, packaging, silicon bring up, characterization, and validation of our broader memory subsystems
  • Independently develop the scripts you need to execute and collaborate with software engineers as your needs scale

A day in the life

A day in the life of a CHDE focused on memory subsystems on the MLA Technology team focuses on operational excellence, constructively identifying problems, prototyping solutions, and leading data collection at scale to improve our products.

We start each day looking at our fleet, reviewing dashboards for emergent issues impacting our customers, partnering with other teams to drive complex debugs as it pertains to UCIe and associated SoC subsystems. We then look forward to the future technologies being developed and how we can best focus our efforts to help improve them and ensure a high quality product on behalf of our customers.

Our team members touch everything from electrical simulations, to hardware qualification on test benches, to software driven data center metrics, with a broad range of tasks across multiple skillsets where you can help improve the reliability and performance of our products.

You help the team evolve by actively participating in design discussions, team planning, code reviews, tickets/metric reviews, and data center capacity initiatives. CHDEs on the MLA Technology Team are expected to help mentor others on the team in their area of expertise to help develop the team’s baseline skillsets and to participate in the hiring process for the team.

  • Bachelor's degree in Electrical Engineering, Computer Engineering, Systems Engineering, or related fields
  • 5+ years of practical semiconductor design work including full-chip and subsystem integration experience
  • Good knowledge of memory training, timing parameters and/or controller features
  • Ability to create scripts (lua, bash, python, etc.) to accomplish functional day to day tasks.
  • Knowledge of DDR/HBM/UCIe phy and controller related protocols
  • Drive cross-functional triage effort on functional and performance issues
  • 3+ years in SOC/IO/Subsystems Experience working closely with physical design teams to develop highly optimized ASICs with excellent power, performance and area
  • Support the physical design team with IP integration, 2.5D packaging, clocking and timing constraints
  • Drive cross-functional triage effort on functional and performance issues
  • Perform system-level debug and root-cause analysis through bring-up, characterization, validation and production phases

Amazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status.

Our inclusive culture empowers Amazonians to deliver the best results for our customers. If you have a disability and need a workplace accommodation or adjustment during the application and hiring process, including support for the interview or onboarding process, please visit https://amazon. jobs/content/en/how-we-hire/accommodations for more information.

If the country/region you’re applying in isn’t listed, please contact your Recruiting Partner.

The base salary range for this position is listed below. Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location.

Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage), 401(k) matching, paid time off, and parental leave. Learn more about our benefits at https://amazon. jobs/en/benefits .

USA, TX, Austin - 159,200. 00 - 215,300. 00 USD annually

Interested in this role?Continue on Amazon's careers page.
Apply on Amazon