Senior / Staff Software Engineer, Product Infrastructure
Spotted 2h agofulltime
AI Agent Apply · Ashby & Greenhouse
Review with AI agent You find the fit. Your agent handles the form.
Choose a role or send your matches to the agent. It uses your original résumé and saved details, applies in the cloud, and keeps every result in one place.
Job description
About this role
Employer-provided description, formatted for easier reading.
About Prometheus
Prometheus is building AI systems for engineering in the physical world. Behind them sits a platform that runs large, long-lived computational workloads: fleets of agents and numerical solvers, very large binary artifacts, and results that have to be correct because real hardware depends on them.
The role
You will build the infrastructure this platform runs on. It is a small set of hard, coupled systems, and you will own one or more of them end to end: design, implementation, operations, and on-call.
- Durable orchestration. Workloads run for days across thousands of tasks. They must survive failure, checkpoint, resume, and stay steerable by humans and agents mid-flight.
- Versioned artifact storage. Git-style semantics (branch, diff, merge) over large binary artifacts, on content-addressed storage with aggressive deduplication, at terabyte-to-petabyte scale.
- CI and validation pipelines. Merge queues where the gates are expensive computed checks, not just unit tests, and a change is validated by actually running it.
- Compute execution fabric. Scheduling heterogeneous workloads, many on GPUs, across cloud fleets: batching, prioritization, budget enforcement, and streaming results back in real time.
- Agent harness. Sandboxed, reproducible, observable execution for long-running autonomous agents that write code, run tools, and spend real compute.
You might be a fit if
- You have 5+ years building and operating distributed systems in production; for the Staff level, 10+ years and a system you are known for.
- You have seen real scale and carry the scars: at a database, streaming, or orchestration company; a multiplayer game backend; internet-scale media storage; a simulation or GPU-compute platform; or a build/CI system serving thousands of engineers.
- You have gone deep in at least one of: durable workflow engines, storage engines or content-addressed storage, high-throughput schedulers, or build/merge systems.
- You write production systems code (Go, Rust, or C++) and are effective in Python.
- You have carried a pager for stateful systems in production.
- You want your work to matter in the physical world.
Strong pluses
- GPU scheduling or HPC job orchestration (Slurm, Ray, Kubernetes batch).
- Deterministic simulation testing or serious correctness culture (FoundationDB, TigerBeetle, Jepsen lineage).
- Exposure to scientific or geometric computing. Not required; we will teach you the domain.
- Experience building infrastructure for AI agents: sandboxing, session persistence, tool execution.
Interested in this role?Continue on Prometheus's careers page.
Apply on Prometheus