Principal Machine Learning Engineer
Spotted 9h agofulltime
Job description
About this role
Employer-provided description, formatted for easier reading.
Join an early-stage startup building the future of AI hardware. Working alongside world-class engineers and leading AI labs, you'll help design workload-specialised accelerators that push the boundaries of AI inference performance.
This is a rare opportunity to work across the full hardware-software stack, shaping how frontier AI models execute on next-generation silicon.
What You'll Do
- Optimise AI inference workloads for custom accelerators
- Improve model performance through quantisation, kernel optimisation and graph-level techniques
- Build performance models to guide hardware design decisions
- Collaborate with compiler, runtime, architecture and hardware teams on hardware-software co-design
- Develop profiling tools and infrastructure to evaluate next-generation AI systems
What You'll Bring
- Strong proficiency in Python, C++, and PyTorch, with a demonstrated history of shipping high- quality software in a startup or fast-paced environment.
- Experience as a developer of one or more LLM inference serving frameworks such as vLLM or SGLang.
- Proficiency in performance analysis required and GPU kernel development (CUDA, Triton, or ROCm) is a plus.
- Comfortable working across software and hardware in a fast-moving start-up environment
Interested in this role?Continue on Oho Group's careers page.
Apply on Oho Group