Software Engineer, Agents (Internal Audit)
You find the fit. Your agent handles the form.
Choose a role or send your matches to the agent. It uses your original résumé and saved details, applies in the cloud, and keeps every result in one place.
What you'll need to apply
Fields this application requires
Company-specific questions
- Describe a time you worked directly with a customer to identify and solve a complex technical problem. What was the problem, how did you approach it, and what was the outcome?essay
- Walk us through a production feature or system you owned end-to-end — from scoping to deployment and iteration. What tradeoffs did you make, and what would you do differently?essay
About this role
Employer-provided description, formatted for easier reading.
About Us
Fieldguide is establishing a new state of trust for global commerce and capital markets by automating and streamlining the work of assurance and audit practitioners — specifically in cybersecurity, privacy, and financial audits. We build software for the people who enable trust between businesses.
We're based in San Francisco, CA, and we're backed by Goldman Sachs Alternatives, Bessemer Venture Partners, 8VC, Floodgate, Y Combinator, and more. Over 50 of the top 100 accounting and consulting firms trust Fieldguide to power mission-critical work.
About the Role
You'll join a genuine 0→1 team on the ground floor of one of the company's biggest new bets. This seat is specifically product-focused : you'll own agent quality, ship agents that do real audit work, and work alongside practitioners.
Depending on your experience and what you're looking to own, you may join building within a major agent area, owning one end-to-end, or setting technical direction for agentic audit work across the team. We're hiring across all levels and will calibrate during interviews based on scope and demonstrated experience.
What You'll Do
- Make agent judgment repeatable: run error analysis on real testing data and turn findings into concrete fixes
- Tradeoffs such as quality/latency/cost across a long multi-phase run
- Build structured-output pipelines that turn model output into real audit artifacts
- Take ambiguous problem statements and turn them into a plan, a shipped feature, and a clear read on what was cut and why
- Work directly with an embedded subject matter expert and with design-partner firms, turning their feedback into agent changes within days
- Expand agent coverage into new controls and new areas of internal audit
Who You Are (All Levels)
- Product-minded and full-stack: you've shipped LLM-backed features to production against real users, and you measure yourself on whether they got used
- You're fluent in evals and error analysis, and you apply them in service of shipping something practitioners trust
- You have real opinions on model selection, prompting, and orchestration tradeoffs, and you can defend them with evidence rather than vibes
- Energized by 0→1 work: you'd rather define the problem than inherit a spec, and you don't stall on ambiguity
- Strong instincts for human-in-the-loop design
- A genuine team player across the organization, not just within engineering: you'll work daily with PM, design, domain experts, and customer-facing teams, and you treat that as the best part of the job
- Ship fast without leaving a mess: your code is reviewable, tested where it counts, and instrumented
- Able to internalize a hard domain fast. You don't need to know SOX today, but you'll understand it well enough to make the right product calls
Higher-Level Responsibilities
At the Senior level, you may:
- Own a major agent area end-to-end, from how the agent reasons about a class of controls through to the artifact a reviewer signs
- Set the evals and error-analysis practice for the team's agent work, and decide what evidence justifies shipping a change or rolling it back
- Collaborate with PMs and designers to shape roadmaps and define architectural tradeoffs, including where the agent acts and where the auditor decides
- Own the harder model and orchestration judgment calls across a long multi-phase run
- Mentor other engineers and raise the bar on 0→1 execution and applied eval rigor
At the Staff level, you may:
- Drive agent initiatives that reach beyond Internal Audit and influence how agents are built across Fieldguide
- Set and champion engineering standards for agent reliability, reproducibility, and defensibility
- Partner with engineering and product leadership to define long-term technical strategy for agentic audit work
- Serve as a trusted advisor to leaders across Engineering, Product, and Design
- Represent Fieldguide externally through writing, speaking, and open-source contributions
Experience
Must-have:
- Shipped LLM-backed product features to production against real users
- Applied AI skillset: evals, error analysis, and model-selection decisions you owned and can explain
- Comfortable full-stack, with enough backend depth to work in agent orchestration
- Autonomy working from an ambiguous spec
- A collaborative mode that works across PM, design, and domain experts
Nice-to-have:
- Python, TypeScript, React, Postgres, Hasura, GraphQL
- Temporal or comparable durable-execution / workflow orchestration
- Hands-on eval experience (Langfuse, Braintrust, LangSmith, Arize Phoenix, or comparable)
- Structured-output work including schema contracts, generating real artifacts from model output
- Startup experience, as a founder or as an early engineer
- A 0→1 track record: things you started where no scaffolding existed
- Experience working directly with customers, and comfort being in the room when they use what you built
- Background in internal audit, SOX, accounting, or another regulated domain
- Document processing, including PDF and Excel manipulation and annotation
Not a fit if:
- Prompt engineering is your whole skill set
- Your agent work never carried production traffic
- You want to own eval methodology or the evaluation harness itself rather than ship product features (better fit on Foundation Agents)
- You need a fully specified ticket to start
- You'd rather not be in the room with customers and domain experts
What Should Excite You
- 0→1 on the biggest bet: You're building the agent and the product from scratch, on the ground floor of where the company is going
- Repeatable judgment: Making an agent reach the same defensible conclusion twice, in a domain where ground truth requires expert judgment
- Real audit stakes: Your work directly affects what firms put in front of their clients, and what a reviewer is willing to sign
- Customer proximity: Design-partner firms and an embedded SOX expert use what you ship within days of it landing
- Human-in-the-loop design: Deciding where the agent acts and where the auditor decides, on work that genuinely matters
- High trust, high autonomy: You're given ambiguous problems and trusted to define the plan
Benefits
- Competitive compensation with equity
- Comprehensive health and wellness benefits
- Flexible time off and work schedules
- Technology reimbursements
- 401(k) plan
- Twice-yearly in-person offsites across the U.S.
- Wellness benefits starting on your first day
Our Values
- Fearless — Inspire and break down seemingly impossible walls
- Fast — Launch fast with excellence; iterate to perfection
- Lovable — Deliver happiness and 11-star experiences
- Owners — Execute and run the business with ownership
- Win-win — Create mutual value and earn trust for life
- Inclusive — Scale the best ideas with inclusive teams