Data Engineer

WorkOS · United States & Canada

Spotted 1h agoFullTime
AI Agent Apply · Ashby & Greenhouse

You find the fit. Your agent handles the form.

Choose a role or send your matches to the agent. It uses your original résumé and saved details, applies in the cloud, and keeps every result in one place.

Review with AI agent

What you'll need to apply

Fields this application requires

NameEmailRésuméPhoneLocationWork authorization answerVisa sponsorship answerLinkedIn profile
Job description

About this role

Employer-provided description, formatted for easier reading.

About WorkOS 🚀

WorkOS builds modern developer tools and APIs that make it easy for companies to become Enterprise Ready. Our platform powers authentication, identity, authorization, and other critical infrastructure that developers need to securely scale their products to large organizations.

We recently raised a $100M Series C, valuing the company at $2B, led by Meritech and Sapphire with participation from Greenoaks, Craft, Abstract, and Audacious. WorkOS powers enterprise features for many of the fastest-growing AI companies, including OpenAI, Cursor, and Perplexity, Sierra, and Plaid.

As AI reshapes software, WorkOS is at the frontier of Human and Agent Authentication, Identity, and Access Control helping companies answer a new critical question: who are your agents, and what are they allowed to do? Our fast-growing customer base includes hundreds of modern software companies building the next generation of enterprise-ready products.

About the role

The Data team at WorkOS plays a central role on the company operations to ensure information and insights are accurate and available from any surface, whether that's a visualization tool or an agent's response, to accelerate company growth and scale. The team owns WorkOS's internal data platform end to end: ingestion, orchestration, the Snowflake warehouse, dbt transformations, governance and access controls.

We also own the consumption layer: the reverse-ETL syncs that land data in Salesforce and Slack, the semantic views that agents query, the definitions that keep reporting consistent across the company, and the visualization tooling the company uses (ask us about Flashboards ).

We run on an agent-first operating system: we are a lean team that moves fast, we document so agents can execute, and AI agents query the warehouse, run our runbooks, and open pull requests alongside us. We trust each other's expertise, we work out loud and in the open, and not opposed to rapid experimentation to get to the right solution for the business.

We're hiring a Data Engineer to own and evolve the systems that move, transform, protect, and serve data across our internal warehouse. This is a high-ownership role on a lean team: you will be the DRI for the platform and its core data models, partner directly with Product Engineering, RevOps, Finance, GTM Engineering, and Security, and raise the bar on reliability, correctness, and operational rigor.

The role spans the platform and what runs on it. You will own ingestion, orchestration, and access governance; the dbt models and metric definitions that depend on them; and how AI is applied across the data stack, so that trusted data is easy for both people and AI agents to discover, understand, and work with.

You don't need to have done this exact job before. The best data engineers at WorkOS are the ones who notice in a Slack thread that a number doesn't match, trace it from the dashboard through the Gold model to the connector, fix the connector, and confirm the definition with its owner before the weekly review.

When a business team asks for a field in Salesforce, they don't hand it off; they build the model, write the sync, and add the test. They would rather automate the runbook than run it a third time. They enjoy collaboration with different parts of the business and can find simple, scalable solutions to ambiguous problems.

If that sounds like how you work, we want to talk.

Responsibilities

  • Own the reliability, freshness, and scaling of the ingestion and orchestration pipelines that land source data in Snowflake, including monitoring, alerting, runbooks, and backfill and reprocessing patterns
  • Design, build, and scale the core dbt models (Bronze, Silver, Gold) for billing and usage, CRM, product events, and GTM funnel reporting
  • Partner with Product, Finance, RevOps, and GTM to define metrics and codify them in the warehouse, and diagnose and resolve data quality and freshness issues at source, not just downstream
  • Own Snowflake RBAC, dynamic masking, and PII classification for humans, agents, and service accounts, so sensitive data is protected without blocking legitimate use, and review DDL and access requests from Engineering and GTM
  • Own the reverse-ETL layer and the runbooks that deliver warehouse data to Salesforce, Slack, and internal agents
  • Extend the semantic views, context, and evaluations that let agents answer business questions accurately, and expand what agents can operate directly, from runbooks and ingestion to pull requests
  • Build and advance the CI/CD review gates in the data-platform monorepo, including the automated review that agent-authored pull requests pass through
  • Own the infrastructure the data platform runs on: the compute, deployments, secrets and access patterns, and environments behind ingestion and orchestration, with attention to availability and failure modes, in partnership with our infrastructure engineers

Example problems you might work on, possible paths rather than a fixed roadmap:

  • Standardizing how Postgres and SaaS sources are ingested, and consolidating end-to-end pipeline deployment and orchestration in Prefect
  • Managing Snowflake RBAC, resource management, and masking policies as code with Terraform
  • Using query history to improve analytics tables and semantic views, and to keep agent answers accurate and consistent as data grows and metrics shift
  • Automating masking coverage as new data sources arrive, so adding a new table no longer needs a human in the loop
  • Automating near-certain matches in the Identity Graph
  • Pseudonymizing product data in the ingestion path, alongside compliance and data-deletion policies
  • Extending transcript aggregation across sources, including PII detection before transcripts enter the warehouse
  • Building the reverse-ETL framework and runbooks that deliver data to other tooling and agents
  • Running self-hosted data services on Kubernetes and managing their AWS resources (IAM, storage, secrets) as code with Terraform

Qualifications

  • 5+ years (or equivalent) building and operating production data platforms, including the transformation layer
  • Deep experience with Snowflake, dbt, and an orchestrator such as Prefect, Airflow, or Dagster, and with ingesting data from production databases and SaaS sources, including backfill and reprocessing patterns
  • Experience running data systems in production: monitoring, alerting, debugging, incident response habits, and pragmatic SLO/SLA thinking
  • Experience with warehouse access governance: RBAC, masking policies, and PII handling
  • Strong data modeling judgment and an understanding of schema evolution
  • Strong SQL and Python; disciplined engineering practices (tests, docs, reviews, CI/CD)
  • Proven experience operating as the only engineer on a layer, balancing foundational architecture with urgent business needs
  • Working knowledge of how GTM, Finance, and Product teams operate in addition to working with Engineering, and experience translating ambiguous business questions into technical specs
  • Comfortable using LLMs and coding agents as part of your day-to-day development workflow, with the judgment to validate their output and scale it into shared processes, and embracing AI and automation to scale platform work and accelerate your teammates
  • A systems thinker who reasons carefully about freshness, correctness, and failure modes, especially for data that everything else depends on
  • Pragmatic: you start with the simplest solution that works, prove it, then scale it, and you balance fast answers with durable solutions
  • You look for where AI can remove a bottleneck for the people and agents who depend on the warehouse, and you build the durable tools and workflows they rely on every day

Benefits and Perks ( US Only) 💖

At WorkOS, we offer resources that emphasize personal and familial well-being. We offer healthcare coverage for you and your family, including medical, dental, and vision. We offer parental leave, paid-time off and fully remote working arrangements.

  • 401k matching
  • Competitive Equity
  • Healthcare, dental and vision coverage
  • FSA, ST/LT Disability, Voluntary Life
  • Carrot fertility benefits
  • 20 days paid vacation + 10 holidays + unlimited sick leave
  • 12 weeks fully paid parental leave
  • Fitness : Monthly stipend for gyms, yoga classes, race registrations or whatever keeps you active
  • Wellness : Monthly stipend for a massage, meditations class, therapy, or activities that enhance your well-being
  • Commuter benefits for hybrid employees in SF/NYC
  • Unlimited token usage!

Please inquire directly with our recruiting team for benefits available to those working outside the US.

Equal Opportunity

Employer

WorkOS is an equal opportunity employer, committed to diversity and inclusiveness. We will consider all qualified applicants without regard to race, color, nationality, gender, gender identity or expression, sexual orientation, religion, disability or age.

We may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans.

If you would like more information about how your data is processed, please contact us.

Interested in this role?Continue on WorkOS's careers page.
Apply on WorkOS