Back to jobs

Lead Infrastructure Engineer

Arrows · New York, United States

Spotted 3d ago

Job details

Pay
$180,000 – $230,000 a year
Work mode
Hybrid
Level
Staff / principal
Experience
4+ years
Posted
Oct 8, 2026
Last confirmed open
Oct 8, 2026
Job description

About this role

About the Company

Founding Infrastructure Engineer New York City, Financial District Hybrid Full-time / YC-backed About the Role

The role

You'll own the platform every one of our agents runs on, the execution layer that turns a 3,000 sheet drawing set into hundreds of verified engineering findings. That means the async worker fleet, the agent sandboxes, the LLM traffic layer, and the document-processing pipeline underneath all of it.

The headline focus is our agent sandboxing platform: sandboxed code execution for LLM agents, warm pools for instant startup, strict isolation, and controlled network egress for untrusted agent generated code. We run on Azure Container Apps and Bicep today. We look for deep expertise in containers, workload isolation, and Azure infrastructure rather than any specific toolchain.

You will lead our infrastructure decisions going forward.

Responsibilities

  • Own the agent sandboxing platform and keep it fast, cheap, and secure as agent traffic grows.
  • Run the async worker fleet that processes thousands of drawing sheets per project, autoscaling from zero to hundreds of jobs and back.
  • Own the LLM traffic layer: keep multi-tenant traffic fair, fast, and metered — every token accounted for in customer billing.
  • Keep streaming reliable: chat and progress updates that survive deploys and disconnects.
  • Scale document processing: rendering, OCR, and extraction pipelines where one project can be tens of gigabytes of drawings.
  • Own CI/CD, reliability, and cost — from the self-hosted runner fleet to incident response to the Azure bill.
  • Ship in the codebase: contribute to our Python/FastAPI backend alongside product engineers.

Qualifications

  • REQUIRED 4+ years building and operating backend or distributed systems in production, including cloud infrastructure (Azure, AWS, GCP); Azure preferred.
  • Required Skills Strong Python, our deploy and environment tooling is Python.
  • Deep production container experience.
  • Azure Container Apps or similar preferred; AKS depth transfers.

Workload isolation and sandboxing

Linux primitives (namespaces, seccomp, bubblewrap/gVisor-class tooling) and safe execution of untrusted code. Expert infrastructure as code in any major tool — we use Bicep today. CI/CD ownership with GitHub Actions or Azure DevOps.

Azure identity and networking

Entra ID, managed identities, RBAC, VNets, private endpoints. Production Postgres and Redis operations.

Observability

OpenTelemetry, Log Analytics/KQL, incident response. Rapid prototyping with AI coding agents (Claude Code or Codex), and the judgment to review and verify what they produce. Preferred Skills NICE TO HAVE Kubernetes experience.

Infrastructure for LLM/ML workloads: inference fleets, agent orchestration, vector search. KEDA or other event-driven autoscaling. Experience with agent frameworks and harnesses (MCP, tool-use loops, Claude Agent SDK, LangChain/LangGraph).

Document processing, computer vision, or AEC industry exposure. Pay range and compensation package Compensation and benefits Base salary $180,000 – $230,000 · Equity up to 1.00% · Health and dental.

Equal Opportunity

Statement YC-backed, well-funded, and already generating revenue. A rare combination: genuinely hard AI and systems problems inside a $2T industry that's barely been touched by software. Huge ownership, early equity, and your work ships to enterprise customers from day one.

Oxford founders and a tight-knit team of 10 engineers who love building things properly. AI agents are part of the engineering workflow, not just the product.

Interested in this role?Continue on LinkedIn to apply.
Apply on LinkedIn