Staff Software Engineer, ML Frameworks
What you'll need to apply
What this employer's standard application typically asks
About this role
Employer-provided description, formatted for easier reading.
Minimum qualifications
- Bachelor’s degree or equivalent practical experience.
- 8 years of experience in software development.
- 5 years of experience testing, and launching software products, and 3 years of experience with software design and architecture.
- 5 years of experience with one or more of the following: Speech/audio (e.g., technology duplicating and responding to the human voice), reinforcement learning (e.g., sequential decision making), ML infrastructure, or specialization in another ML field.
- 5 years of experience with ML design and ML infrastructure (e.g., model deployment, model evaluation, data processing, debugging, fine tuning).
Preferred qualifications
- Master’s degree or PhD in Engineering, Computer Science, or a related technical field.
- 8 years of experience with data structures and algorithms.
- 3 years of experience in a technical leadership role leading project teams and setting technical direction.
- 3 years of experience working in a complex, matrixed organization involving cross-functional, or cross-business projects.
- Experience in ML compilers and runtimes.
- Experience with Tensor Processing Units (TPUs), TPU system design, and Graphics Processing Units (GPUs).
About the job
Google's software engineers develop the next-generation technologies that change how billions of users connect, explore, and interact with information and one another. Our products need to handle information at massive scale, and extend well beyond web search.
We're looking for engineers who bring fresh ideas from all areas, including information retrieval, distributed computing, large-scale system design, networking and data storage, security, artificial intelligence, natural language processing, UI design and mobile; the list goes on and is growing every day.
As a software engineer, you will work on a specific project critical to Google’s needs with opportunities to switch teams and projects as you and our fast-paced business grow and evolve. We need our engineers to be versatile, display leadership qualities and be enthusiastic to take on new problems across the full-stack as we continue to push technology forward.
With your technical expertise you will manage project priorities, deadlines, and deliverables. You will design, develop, test, deploy, maintain, and enhance software solutions.
Eliminate resource waste across Google’s accelerator fleet, maximizing the physical utility of compute clusters while maintaining peak developer velocity and seamless runtime execution. We strive to provide a cohesive, highly efficient, and transparent runtime environment that enables ML teams to focus entirely on modeling and research rather than physical infrastructure constraints.
Turn Google’s compute infrastructure into a completely fluid and self-optimizing accelerator ecosystem. We envision a future where both internal product teams and external Google Cloud/hybrid-cloud customers can access and share massive accelerator pools on-demand with zero cold-start latency, zero wasted idle capacity.
Google Cloud accelerates every organization’s ability to digitally transform its business and industry. We deliver enterprise-grade solutions that leverage Google’s cutting-edge technology, and tools that help developers build more sustainably. Customers in more than 200 countries and territories turn to Google Cloud as their trusted partner to enable growth and solve their most critical business problems.
Individual pay is determined by factors including job-related skills, experience, and relevant education or training.
US: $207000 - $300000 (USD) + 20% bonus target + equity + benefits
Learn more about benefits at Google .
Responsibilities
- Drive technical strategy, roadmaps, and adoption for large-scale ML infrastructure development.
- Exercise sound engineering judgment to guide sustainable engineering choices for ML systems at scale.
- Seek additional opportunities to drive efficiencies in ML workloads using scaling, idle suspend, and improving these capabilities with existing and novel technologies.
- Deliver impactful software features for Google's accelerator fleet.
- Partner with GDM and other PAs to transfer key innovations into products developed by your team and partner teams.