Director, Site Reliability Engineering - Paze

Early Warning Services · Scottsdale, AZ, US

Spotted 20h agofulltime
Job description

About this role

Employer-provided description, formatted for easier reading.

At Early Warning, we’ve powered and protected the U.S. financial system for over thirty years with cutting-edge solutions like Zelle®, Paze®, and so much more. As a trusted name in payments, we partner with thousands of institutions to increase access to financial services and protect transactions for hundreds of millions of consumers and small businesses.

Early Warning follows a hybrid work model to allow for a more collaborative working environment.

Candidates responding to this posting must independently possess the eligibility to work in the United States, for any employer, at the date of hire. This position is ineligible for employment Visa sponsorship.

Role Summary

The Director, Site Reliability Engineering leads the SRE function for an assigned product, platform, pillar or business domain and is accountable for the reliability, scalability, performance and operability of its business-critical production services.

The Director develops engineering talent, establishes domain reliability strategy and priorities, and partners with Product, Software Engineering, Platform, Infrastructure, Security and Risk teams relevant to the domain.

The role translates business and product priorities into measurable reliability outcomes and ensures teams have the engineering practices, capabilities and operating mechanisms needed to achieve them.Core SRE Responsibilities

=============================

1. Reliability Objectives and Measurement

----------------------------------------------

  • Establish meaningful Service Level Indicators (SLIs) and Service Level Objectives (SLOs) aligned with customer, product and business outcomes.
  • Use quantitative production data to measure performance and partner with domain teams to set reliability targets appropriate to business criticality, architecture, customer expectations and cost.

2. Reliability Risk and Error Budgets

------------------------------------------

  • Use SLO performance, error budgets, failure data and production evidence to identify and prioritize reliability risk.
  • Ensure persistent risks have accountable engineering plans, appropriate escalation and sustained follow-through.

3. Software Engineering and Automation

----------------------------------------------

  • Champion reusable software, tooling and automation that improve reliability, scalability and operability while reducing repetitive operational work.
  • Ensure SRE-developed solutions follow sound engineering practices, including source control, code review, testing, maintainability and secure development.
  • The Director is expected to remain hands-on and maintain sufficient technical depth to develop code, contribute to automation and engineering solutions, and work directly with engineers when appropriate.

4. Observability Engineering

------------------------------------

  • Establish expectations for metrics, logs, traces, dashboards and telemetry sufficient to understand service behavior, customer impact and systemic risk.
  • Drive actionable alerting and observability that support rapid diagnosis, performance analysis, capacity planning and continuous improvement.

5. Incident Management and Service Restoration

---------------------------------------------------

  • Provide domain leadership and accountability for disciplined incident response, service restoration and stakeholder communication.
  • Improve response through automation, runbooks, training, exercises and analysis of recurring patterns.

6. Learning From Failure

-----------------------------

  • Promote blameless post-incident review focused on systemic learning, prevention and measurable follow-through.
  • Use patterns across incidents and near misses to drive architectural, operational and domain-level improvements.

7. Production Readiness, Resilience and Capacity

-----------------------------------------------------

  • Ensure critical services meet appropriate production-readiness, resilience, recovery, capacity and scalability expectations.
  • Partner with engineering teams to address systemic failure modes, dependency risks, capacity constraints and recovery gaps before they affect customers.

8. Operational Toil and Sustainable Engineering

----------------------------------------------------

  • Measure and reduce repetitive, manual and low-value operational work through engineering, automation and simplification.
  • Ensure operational responsibilities inform engineering priorities without becoming the primary definition of the SRE role or creating unsustainable team load.

9. Security, Risk and Compliance

-------------------------------------

  • Partner with Security, Risk, Compliance and engineering teams relevant to the domain to meet EWS control, resilience and regulatory obligations.
  • Ensure material reliability risks and control gaps are visible and addressed within the domain or escalated when they exceed its authority.

The Director operates within an assigned domain and delivers outcomes through its SRE teams. Impact is demonstrated by building strong teams and leaders, setting direction, resolving barriers within the domain and with direct dependencies, and creating durable engineering mechanisms. The Director leads through others while remaining sufficiently hands-on to contribute directly when appropriate.

Success is measured primarily by the reliability outcomes and capabilities of the teams the Director leads.

Leadership Responsibilities

===============================

  • Own SRE strategy, execution and measurable reliability outcomes for the assigned domain.
  • Build, develop and retain high-performing SRE teams with clear accountability, career development and succession.
  • Translate business and product priorities into reliability investments and an executable multi-quarter roadmap.
  • Set priorities and make evidence-based tradeoffs across reliability, delivery, operational risk, capacity and business needs.
  • Develop capable managers and technical leaders, ensuring decisions are made at the appropriate level.
  • Partner with engineering, product, infrastructure, security and risk teams relevant to the domain to embed reliability into engineering decisions.
  • Make reliability risk, SLO performance, operational load and improvement progress visible through effective metrics and operating reviews.
  • Provide accountable leadership during significant incidents and ensure systemic corrective actions are completed.

Required Qualifications

===========================

  • Typically 12 + years of relevant software engineering, site reliability engineering, production engineering, platform engineering or closely related experience, including significant technical leadership.
  • 5+ years of people leadership experience, with demonstrated success leading engineering teams and developing managers and/or senior technical leaders.
  • Experience operating highly available, business-critical distributed systems and leading reliability improvement across a major product, platform, pillar or business domain.
  • Strong understanding of SRE practices, including SLOs/SLIs, error budgets, observability, incident management, automation, capacity, resilience and production readiness.
  • Demonstrated hands-on technical capability and sufficient depth to develop code, contribute to automation and engineering solutions, and work directly with engineers when appropriate.
  • Ability to communicate technical risk, tradeoffs and investment needs clearly to engineering, product and senior business stakeholders.
  • Demonstrated ability to build inclusive, accountable and high-performing engineering teams.

Preferred Qualifications

============================

  • Experience in payments, financial services or another highly regulated, high-availability environment.
  • Experience leading SRE or production engineering across multiple teams or a complex product or platform ecosystem.
  • Experience with cloud platforms, distributed systems, modern observability, infrastructure automation and software delivery at scale.
  • Experience establishing reliability metrics, governance and operating reviews across teams within a defined domain.

The base pay scale for this position in

Phoenix, AZ/ Chicago, IL in USD per year is: $173,000 - $230,000.

San Francisco, CA in USD per year is: $207,000 - $276,000.

Additionally, candidates are eligible for a discretionary incentive plan and benefits.

Some of the Ways We Prioritize Your Health and Happiness

  • Healthcare Coverage – Competitive medical (PPO/HDHP), dental, and vision plans as well as company contributions to your Health Savings Account (HSA) or pre-tax savings through flexible spending accounts (FSA) for commuting, health & dependent care expenses.
  • 401(k) Retirement Plan – Featuring a 100% Company Safe Harbor Match on your first 6% deferral immediately upon eligibility.
  • Paid Time Off – Flexible Time Off for Exempt (salaried) employees, as well as generous PTO for Non-Exempt (hourly) employees, plus 11 paid company holidays and a paid volunteer day.
  • 12 weeks of Paid Parental Leave
  • Maven Family Planning – provides support through your Parenting journey including egg freezing, fertility, adoption, surrogacy, pregnancy, postpartum, early pediatrics, and returning to work.

And SO much more! We continue to enhance our program, so be sure to check our Benefits page here for the latest. Our team can share more during the interview process!

Pursuant to the San Francisco Fair Chance Ordinance, we will consider for employment qualified applicants with arrest and conviction records.

Early Warning Services, LLC (“Early Warning”) considers for employment, hires, retains and promotes qualified candidates on the basis of ability, potential, and valid qualifications without regard to race, religious creed, religion, color, sex, sexual orientation, genetic information, gender, gender identity, gender expression, age, national origin, ancestry, citizenship, protected veteran or disability status or any factor prohibited by law, and as such affirms in policy and practice to support and promote equal employment opportunity and affirmative action, in accordance with all applicable federal, state, and municipal laws.

The company also prohibits discrimination on other bases such as medical condition, marital status or any other factor that is irrelevant to the performance of our employees.

Interested in this role?Continue on Early Warning Services's careers page.
Apply on Early Warning Services