Software Engineer, Core Infrastructure

Radical AI · New York, NY

Spotted 1h agofulltime
Job description

About this role

Employer-provided description, formatted for easier reading.

The opportunity

We're looking for a senior Software Engineer to own the core infrastructure behind our self-driving lab: the execution planner that orchestrates the physical lab, the backend services and APIs that manage its data, and the compute infrastructure that powers our agents and scientific simulations. The software you'll own moves real samples through a working lab, so the bar for correctness and reliability is high.

You'll report to the Director of Software Engineering and work closely with the rest of the software engineering team, the roboticists, and hardware engineers who build the lab. As a senior engineer, you'll set the technical direction for these systems and lead the projects that change them.

What you'll do

Lab orchestration

  • Own the execution planner that schedules every sample's path through the lab. It resolves conflicts and batches samples across actions that range from a two-second arm move to a multi-day anneal
  • Own the service that decides which software can use which instrument, and move it to durable execution so restarts and upgrades don't interrupt experiments
  • Build simulation and debugging tools to test planner changes and replay problems from production

Backend services and APIs

  • Own the Go services and gRPC APIs that track every sample, experiment and campaign, and the services that run simulation and training jobs on our clusters
  • Prepare the lab backend to run more than one lab and to deploy without downtime

Application and compute infrastructure

  • Run our Kubernetes clusters on AWS, at Voltage Park and in our lab, defined in Pulumi and deployed with ArgoCD
  • Schedule shared GPUs so that simulation and ML training jobs from every team keep them busy

Required qualifications

Core backend and infrastructure expertise

  • At least 7 years of professional experience building and operating production backend systems
  • Strong server-side experience in Go or Rust
  • Experience building fault-tolerant distributed systems that handle concurrency, timeouts, retries and safe rollback
  • Experience running Kubernetes in production with infrastructure as code (Pulumi preferred), and comfort debugging Linux hosts and networks
  • Working knowledge of gRPC and Protocol Buffers, and of MongoDB or another document store

Communication and collaboration

  • Clear, precise communicator, especially when working across disciplines
  • You've led a large technical project from design to production, and you go learn what you don't know
  • You take ownership and find problems before they find users
  • We actively use AI-assisted development and expect engineers to find their own productive workflows with these tools

Pluses

  • Scheduling and planning algorithms, such as multi-agent path finding
  • Experience with Restate for durable execution or Ray for distributed compute
  • Experience running storage and observability on bare-metal Kubernetes, with tools such as Longhorn, OpenTelemetry, Prometheus and Grafana
  • Software that controls physical equipment, in lab automation, manufacturing or robotics
  • Early-stage startup experience

What we offer

  • Work closely with a team on the cutting edge of AI research
  • A mission: an opportunity to fundamentally change the way humanity makes progress through materials science discovery

Compensation

The base salary range for this role is $210,000–$295,000. Where an offer lands depends on experience, skills, and scope.

Equal opportunity

Radical AI is committed to equal employment opportunities regardless of race, color, ancestry, national origin, religion, sex, age, sexual orientation, gender identity and expression, marital status, disability, or veteran status.

Interested in this role?Continue on Radical AI's careers page.
Apply on Radical AI