Principal Machine Learning Engineer

Oho Group · San Francisco, CA

Spotted 9h agofulltime
Job description

About this role

Employer-provided description, formatted for easier reading.

Join an early-stage startup building the future of AI hardware. Working alongside world-class engineers and leading AI labs, you'll help design workload-specialised accelerators that push the boundaries of AI inference performance.

This is a rare opportunity to work across the full hardware-software stack, shaping how frontier AI models execute on next-generation silicon.

What You'll Do

  • Optimise AI inference workloads for custom accelerators
  • Improve model performance through quantisation, kernel optimisation and graph-level techniques
  • Build performance models to guide hardware design decisions
  • Collaborate with compiler, runtime, architecture and hardware teams on hardware-software co-design
  • Develop profiling tools and infrastructure to evaluate next-generation AI systems

What You'll Bring

  • Strong proficiency in Python, C++, and PyTorch, with a demonstrated history of shipping high- quality software in a startup or fast-paced environment.
  • Experience as a developer of one or more LLM inference serving frameworks such as vLLM or SGLang.
  • Proficiency in performance analysis required and GPU kernel development (CUDA, Triton, or ROCm) is a plus.
  • Comfortable working across software and hardware in a fast-moving start-up environment
Interested in this role?Continue on Oho Group's careers page.
Apply on Oho Group