Back to jobs

Lead Databricks AI Engineer - Direct hire only

InstantServe Healthcare · Cary, NC

Spotted 3d ago

Job details

Pay
$140,000 – $145,000 a year
Work mode
Hybrid
Level
Staff / principal
Experience
12+ years
Posted
Oct 8, 2026
Last confirmed open
Oct 8, 2026
Job description

About this role

Lead Databricks AI Engineer

Location: Cary, NC (On-site / Hybrid)

Experience: 12–18 years

Employment: Full-Time

Salary: $140K–$145K per annum + benefits

Eligibility: Due to work Nature US Citizenship required

Must-Have Skills & Experience

Expert-level Python, Scala, and PySpark: production-ready, modular, well-tested solutions; Spark workload troubleshooting; optimization of large-scale batch and streaming pipelines using Delta Lake.

Strong SQL and data modeling (dimensional and normalized), schema design, and data contracts.

Deep Databricks expertise

Delta Lake, Unity Catalog, Jobs & Workflows, cluster and pool management, performance tuning, and Model Serving.

Solid experience with the Azure data stack

  • ADLS Gen2 (zone design, ACLs, lifecycle), Azure Data Factory (parameterized/metadata-driven frameworks), Azure Event Hubs/Kafka, and related services.
  • 3+ years designing and shipping LLM-based systems in production: RAG pipelines, agentic/tool-calling workflows, chunking and embedding strategy, vector and hybrid retrieval, prompt engineering.
  • Evaluation discipline: golden datasets, regression suites, accuracy and hallucination tracking, human-in-the-loop feedback.
  • Hands-on with LangChain, LlamaIndex, or LangGraph, plus at least one provider stack (Azure OpenAI, OpenAI, or Databricks Model Serving).
  • Experience with metadata-driven frameworks: schema inference, data profiling, lineage, and catalogs.
  • 12–18 years of total experience in data engineering / data platform delivery.
  • Proven enterprise-scale delivery of a medallion / lakehouse architecture.

Strong grasp of Azure security and governance

Entra ID, managed identities, RBAC, POSIX ACLs, Key Vault, private endpoints, PII handling.

CI/CD and IaC

Azure DevOps, Terraform, Databricks Asset Bundles, automated testing of data pipelines.

Clear technical writing and the ability to present and defend designs to both engineers and non-technical stakeholders.

Strongly Preferred

Knowledge graphs and ontologies (RDF/SPARQL, Neo4j, graph modeling over a lakehouse).

Text-to-SQL or semantic-layer-backed natural-language query systems at enterprise scale.

ML-based anomaly detection on time-series or transactional financial data.

Financial services or insurance domain experience (finance close, GL, subledger, reconciliation, actuarial data).

LLMOps / MLOps: model and prompt versioning, cost governance, observability.

Certifications

Databricks Data Engineer Professional, Azure DP-203 / DP-700, or AZ-305.

Experience with dbt, Great Expectations, or similar; exposure to Workday, Prism, or Accounting Center.

What This Role Owns

Data Platform Architecture & Engineering

Design and implement lakehouse architecture: Bronze/Silver/Gold contracts, ADLS Gen2 zone layout, Delta Lake table design, partitioning, schema evolution, and retention policies.

Build metadata-driven, parameterized ingestion frameworks for batch files, database extracts, CDC feeds, and streaming (Azure Event Hubs / Kafka, Spark Structured Streaming).

Develop canonical PySpark and Scala Spark jobs; define coding and testing standards; conduct PR reviews; debug production incidents; and tune Spark clusters with cost guardrails.

Implement CI/CD for Databricks and ADF in Azure DevOps using Databricks Asset Bundles and Terraform; establish observability with Azure Monitor and Log Analytics.

AI-Augmented Ingestion & Canonical Mapping

Enable auto-generated bridge documents, DML, and canonical table definitions.

Build AI-assisted source-to-canonical mapping with human review gates.

Implement AI-driven data quality and anomaly detection (data drift, schema drift, volume shifts, reconciliation breaks), automated reconciliation, and synthetic privacy-preserving test data.

Semantic Layer, Knowledge Graph & Conversational Access

Design and deliver a semantic layer and knowledge graph over the lakehouse.

Build a GPT-powered conversational interface (text-to-SQL / semantic-layer retrieval) with row- and column-level security.

Governance & Technical Leadership

Govern data and AI using Unity Catalog (lineage, access control, PII standards).

Present and defend designs in Architecture Review Boards and AI governance forums.

Mentor engineers, establish best practices, and produce clear technical documentation.

Interested in this role?Continue on LinkedIn to apply.
Apply on LinkedIn