Lead Databricks AI Engineer - Direct hire only
Job details
- Pay
- $140,000 – $145,000 a year
- Work mode
- Hybrid
- Level
- Staff / principal
- Experience
- 12+ years
- Posted
- Oct 8, 2026
- Last confirmed open
- Oct 8, 2026
About this role
Lead Databricks AI Engineer
Location: Cary, NC (On-site / Hybrid)
Experience: 12–18 years
Employment: Full-Time
Salary: $140K–$145K per annum + benefits
Eligibility: Due to work Nature US Citizenship required
Must-Have Skills & Experience
Expert-level Python, Scala, and PySpark: production-ready, modular, well-tested solutions; Spark workload troubleshooting; optimization of large-scale batch and streaming pipelines using Delta Lake.
Strong SQL and data modeling (dimensional and normalized), schema design, and data contracts.
Deep Databricks expertise
Delta Lake, Unity Catalog, Jobs & Workflows, cluster and pool management, performance tuning, and Model Serving.
Solid experience with the Azure data stack
- ADLS Gen2 (zone design, ACLs, lifecycle), Azure Data Factory (parameterized/metadata-driven frameworks), Azure Event Hubs/Kafka, and related services.
- 3+ years designing and shipping LLM-based systems in production: RAG pipelines, agentic/tool-calling workflows, chunking and embedding strategy, vector and hybrid retrieval, prompt engineering.
- Evaluation discipline: golden datasets, regression suites, accuracy and hallucination tracking, human-in-the-loop feedback.
- Hands-on with LangChain, LlamaIndex, or LangGraph, plus at least one provider stack (Azure OpenAI, OpenAI, or Databricks Model Serving).
- Experience with metadata-driven frameworks: schema inference, data profiling, lineage, and catalogs.
- 12–18 years of total experience in data engineering / data platform delivery.
- Proven enterprise-scale delivery of a medallion / lakehouse architecture.
Strong grasp of Azure security and governance
Entra ID, managed identities, RBAC, POSIX ACLs, Key Vault, private endpoints, PII handling.
CI/CD and IaC
Azure DevOps, Terraform, Databricks Asset Bundles, automated testing of data pipelines.
Clear technical writing and the ability to present and defend designs to both engineers and non-technical stakeholders.
Strongly Preferred
Knowledge graphs and ontologies (RDF/SPARQL, Neo4j, graph modeling over a lakehouse).
Text-to-SQL or semantic-layer-backed natural-language query systems at enterprise scale.
ML-based anomaly detection on time-series or transactional financial data.
Financial services or insurance domain experience (finance close, GL, subledger, reconciliation, actuarial data).
LLMOps / MLOps: model and prompt versioning, cost governance, observability.
Certifications
Databricks Data Engineer Professional, Azure DP-203 / DP-700, or AZ-305.
Experience with dbt, Great Expectations, or similar; exposure to Workday, Prism, or Accounting Center.
What This Role Owns
Data Platform Architecture & Engineering
Design and implement lakehouse architecture: Bronze/Silver/Gold contracts, ADLS Gen2 zone layout, Delta Lake table design, partitioning, schema evolution, and retention policies.
Build metadata-driven, parameterized ingestion frameworks for batch files, database extracts, CDC feeds, and streaming (Azure Event Hubs / Kafka, Spark Structured Streaming).
Develop canonical PySpark and Scala Spark jobs; define coding and testing standards; conduct PR reviews; debug production incidents; and tune Spark clusters with cost guardrails.
Implement CI/CD for Databricks and ADF in Azure DevOps using Databricks Asset Bundles and Terraform; establish observability with Azure Monitor and Log Analytics.
AI-Augmented Ingestion & Canonical Mapping
Enable auto-generated bridge documents, DML, and canonical table definitions.
Build AI-assisted source-to-canonical mapping with human review gates.
Implement AI-driven data quality and anomaly detection (data drift, schema drift, volume shifts, reconciliation breaks), automated reconciliation, and synthetic privacy-preserving test data.
Semantic Layer, Knowledge Graph & Conversational Access
Design and deliver a semantic layer and knowledge graph over the lakehouse.
Build a GPT-powered conversational interface (text-to-SQL / semantic-layer retrieval) with row- and column-level security.
Governance & Technical Leadership
Govern data and AI using Unity Catalog (lineage, access control, PII standards).
Present and defend designs in Architecture Review Boards and AI governance forums.
Mentor engineers, establish best practices, and produce clear technical documentation.