Software Engineer, Systems ML Tooling

Meta · Menlo Park, CA

Spotted 2h ago

What you'll need to apply

What this employer's standard application typically asks

NameEmailPhoneRésuméLocationWork authorization answerVisa sponsorship answerHow you heard about them
Job description

About this role

Employer-provided description, formatted for easier reading.

Meta is seeking a Software Engineer to join the MTIA (Meta Training & Inference Accelerator) Software Tooling team, which develops and maintains the tooling ecosystem for Meta's in-house AI accelerator ASICs.

The Tooling team provides debugging, profiling, memory analysis, and monitoring capabilities for the whole MTIA Ecosystem, redefining ML accelerator tooling by leveraging Meta's full-stack ownership from silicon specs to fleet observability.

In this role, you will design and build developer tools across the MTIA tooling ecosystem, with a primary focus on debugging, sanitizer technology, and fault isolation for AI workloads running on MTIA hardware at scale.

You will work at the intersection of compilers, runtime, hardware, and ML frameworks, collaborating with cross-functional partners to deliver a high-quality developer experience for Meta's custom AI accelerators.

This role can focus on either debugging and sanitizing technology or simulation infrastructure for pre-silicon and post-silicon model bring-up, depending on the candidate's background and team needs.

Responsibilities

  • Lead the design and development of MTIA's debugging and sanitizer tooling, including live debugging, core dump, and memory sanitizers for accelerator workloads
  • Build debugging capabilities spanning graph-mode debugging, kernel-level diagnostics, and multi-rank fault isolation
  • Contribute broadly to the MTIA SW tooling infrastructure — profiling, performance debugging, memory analysis, monitoring, and reliability analysis for training and inference workloads
  • Own significant components end-to-end, from technical design through implementation, production rollout, and long-term maintenance
  • Collaborate closely with the MTIA compiler, runtime, kernel, and hardware teams to instrument the software stack with the hooks, sanitizers, and debuggers required
  • Drive AI-native tooling approaches, leveraging automation and LLM-guided diagnostics to reduce time-to-root-cause
  • Partner with internal product teams across advertising, recommendations, and generative AI to understand developer pain points and prioritize tooling investments
  • Communicate architectural decisions through design documents and cross-team reviews; share debugging methodology and best practices with engineers across the MTIA stack

Qualifications

  • Bachelor's degree in Computer Science, Computer Engineering, relevant technical field, or equivalent practical experience
  • 4+ years of experience in systems software engineering, performance engineering, developer tooling, or a closely related field
  • Experience building debugging, sanitizer, profiling, simulation, or diagnostic tools for complex software/hardware systems
  • Proficiency in C++ and Python, including low-level systems programming and scripting for tool automation
  • Experience working across multiple layers of a system stack (compiler, runtime, OS/driver, hardware)Track record of leading the technical design and delivery of tooling or infrastructure projects through to production deployment
  • Cross-stack debugging skills, with the ability to trace issues across application, runtime, and hardware boundaries 5+ years of experience in systems software, developer tooling, or accelerator software development (or equivalent with advanced degree)
  • Experience with compiler-based instrumentation and runtime shadow-memory techniques underlying sanitizers (LLVM instrumentation passes, interceptors, allocator red zones)
  • Track record of building developer tools adopted by large engineering populations, ideally with contributions to open-source projects (LLVM sanitizers, gdb, Valgrind, Triton, etc.)
  • Familiarity with ML framework internals (PyTorch graph execution, torch.compile, operator dispatch) and AI compiler stacks (MLIR, LLVM, TVM, Triton)
  • Experience with accelerator ecosystems (GPU/CUDA, TPU, custom ASICs), including memory analysis, performance profiling, and runtime debugging using their toolchains (cuda-gdb, nsight-compute, nsight-systems)
  • Experience working at the device-software boundary — accelerator runtime/driver internals, DMA and memory-mapped device behavior, on-device fault handling or firmware-assisted error reporting
  • Demonstrated ability to integrate AI tools to optimize/redesign workflows and drive measurable impact (e.g., efficiency gains, quality improvements)
  • Experience with Linux debugging and profiling infrastructure (gdb, perf, eBPF, ftrace, coredump analysis, hardware performance counters) and familiarity with binary formats and debugging metadata (ELF/DWARF)
  • Experience with Linux kernel and driver-level debugging, hands-on experience with sanitizer and dynamic-analysis technology, or building comparable memory checking tools
  • Experience with functional or cycle-approximate simulation techniques — instruction-set interpreters/simulators, timing and performance models — and with simulator or virtual-platform frameworks (gem5, QEMU, AModel, BModel)
  • Demonstrated ongoing AI skill development (e.g., prompt/context engineering, agent orchestration) and staying current with emerging AI technologies
  • Experience adhering to and implementing responsible, ethical AI practices (e.g., risk assessment, bias mitigation, quality and accuracy reviews)
Interested in this role?Continue on Meta's careers page.
Apply on Meta