Observability - Site Reliability Engineer
Job details
- Pay
- $112,000 – $179,000 a year
- Work mode
- On-site
- Level
- Senior
- Experience
- 5+ years
- Education
- Bachelor's degree
- Posted
- Oct 11, 2026
- Last confirmed open
- Oct 11, 2026
About this role
About Peraton
Peraton is a next-generation national security company that drives missions of consequence spanning the globe and extending to the farthest reaches of the galaxy. As the world’s leading mission capability integrator and transformative enterprise IT provider, we deliver trusted, highly differentiated solutions and technologies to protect our nation and allies.
Peraton operates at the critical nexus between traditional and nontraditional threats across all domains: land, sea, space, air, and cyberspace. The company serves as a valued partner to essential government agencies and supports every branch of the U.S. armed forces. Each day, our employees solve the most daunting challenges that our customers face.
Visit peraton.com to learn how we’re keeping people around the world safe and secure.
About The Role
We are seeking Observability / Site Reliability Engineers to design, implement, and operate telemetry and reliability capabilities across enterprise infrastructure, platform services, applications, and workloads.
The engineers will implement centralized metrics, logs, traces, dashboards, alerts, and reliability practices using OpenTelemetry or equivalent industry standards and approved enterprise tooling. This role combines observability engineering, site reliability practices, production operations, automation, and performance analysis to improve system health, availability, and operational resilience.
Responsibilities
Telemetry Engineering
- Implement standardized collection of metrics, logs, and distributed traces across enterprise environments.
- Establish OpenTelemetry-based instrumentation and telemetry collection patterns where appropriate.
- Integrate cloud, Kubernetes, platform, application, and infrastructure telemetry into centralized observability services.
- Maintain telemetry pipelines, retention, routing, and data-quality controls.
- Develop and maintain instrumentation and telemetry configurations to support effective system monitoring and troubleshooting.
Reliability & Performance Engineering
- Develop dashboards, alerts, service-health views, and operational metrics to provide visibility into system performance and availability.
- Define and monitor Service Level Objectives (SLOs), availability indicators, latency, error rates, saturation, and capacity measures.
- Analyze system performance, capacity, utilization, and operational trends to identify potential reliability issues and corrective actions.
- Apply site reliability engineering practices to improve system availability, performance, scalability, and operational resilience.
- Identify opportunities to improve detection, response, and recovery across enterprise services.
Operations & Automation
- Support incident diagnosis, cross-service troubleshooting, and root-cause analysis for complex reliability and observability issues.
- Automate recurring monitoring, alerting, reporting, and remediation activities where appropriate.
- Partner with operations and engineering teams to improve incident detection and response and reduce recurring issues.
- Participate in Tier 3/4 escalation and on-call support for observability, platform, and reliability issues.
- Develop and maintain operational procedures, troubleshooting documentation, and knowledge articles.
- Support continuous improvement of monitoring and reliability practices based on operational data and incident trends.
Location
This position is fully onsite in Chantilly, VA
Qualifications
Required Qualifications
- Active TS/SCI clearance with CI Polygraph.
- Approximately 5–8 years of experience in observability, site reliability engineering, operations, platform engineering, cloud engineering, or a related technical discipline.
- 5 years experience with a BS/BA, an additional 4 years of experience may be considered in lieu of a degree.
- Hands-on experience with metrics, logging, tracing, monitoring, and alerting.
- Experience with Kubernetes and cloud environments.
- Strong troubleshooting, systems analysis, and performance-analysis skills.
- Experience operating and supporting production services and responding to technical incidents.
- Experience with scripting, automation, or infrastructure-as-code practices.
- Strong analytical, problem-solving, and communication skills.
Desired Qualifications
Experience with one or more of the following:
- OpenTelemetry.
- Prometheus and Grafana.
- CloudWatch, OpenSearch, or comparable cloud monitoring and observability platforms.
- Distributed tracing and application performance monitoring.
- SLO, SLA, and error-budget practices.
- Observability for large-scale or distributed enterprise platforms.
- Automated incident response or remediation.
- Observability within secure or highly regulated enterprise environment
Details
Target Salary Range: $112,000 - $179,000. This represents the typical salary range for this position. Salary is determined by various factors, including but not limited to, the scope and responsibilities of the position, the individual’s experience, education, knowledge, skills, and competencies, as well as geographic location and business and contract considerations.
Depending on the position, employees may be eligible for overtime, shift differential, and a discretionary bonus in addition to base pay.
Benefits
- Statement: Peraton offers eligible employees a variety of benefits including medical, dental, vision, life, health savings account, short/long term disability, EAP, parental leave, 401(k), paid time off (PTO) for vacation, and company paid holidays.
- A full listing of available benefits can be viewed at https://www.careers.peraton.com/benefits.
Application Statements
The application period for the job is estimated to be 30 days from the job posting date. However, this timeline may be shortened or extended depending on business needs and the availability of qualified candidates. By applying to this job, you are expressing interest in the role and the Company.
During the review of your application, you may be required to participate in an on-camera interview, as well as participate in a process to verify your identity. Use of artificial intelligence (AI) tools of any kind during Peraton interviews is strictly prohibited unless the candidate has obtained prior written authorization. All interview responses must be the candidate’s own.
EEO: Equal opportunity employer, including disability and protected veterans, or other characteristics protected by law.