Sr. Platform Engineer
Job details
- Pay
- $150,000 – $170,000 a year
- Employment
- Full-time
- Level
- Senior
- Experience
- 4+ years
- Posted
- Oct 9, 2026
- Last confirmed open
- Oct 10, 2026
About this role
Job Title: Sr. Platform Engineer
Location-Type: Remote
Start Date: ASAP
Duration: Permanent
Compensation Range: $150,000 - $170,000
Benefits: Medical, Telehealth, Dental, Vision, 401(k), Health Savings Accounts (HSA), Flexible Spending Accounts (FSA), Life and AD&D, Short-Term and Long-Term Disability, Flex PTO, Leave of Absence, Employee Assistance Program, Wellness Program, Rewards and Recognition Program; 3% annual bonus
Visa Sponsorship: Not eligible for visa sponsorship
Job Description:
This hands-on platform engineering role is responsible for building and operating the client's core IT platforms, including observability, DevOps, ITSM, and integrations, with direct ownership over critical technology roadmap outcomes.
Job Summary
•Design, develop, and manage automated, resilient, self-healing platforms with native-AI capabilities serving internal and customer business needs.
•Build and manage the observability backend stack, including Grafana Enterprise, Mimir, Loki, Tempo, and Alertmanager on Kubernetes via Helm and GitLab CI/CD.
•Author and maintain infrastructure-as-code using Terraform modules, Ansible AWX playbooks, ArgoCD, and Helm for automated provisioning and deployment.
•Develop and expand the OpenTelemetry Prometheus scrape profile library, including SNMP, REST API, and cloud provider exporters across multiple device classes.
•Build AIOps capabilities including anomaly detection integrations, event correlation rules, and synthetic monitoring to reduce alert noise.
•Integrate Alertmanager with ITSM tooling, including webhook routing, ticket enrichment, auto-close logic, and escalation policy configuration.
•Mentor mid-level engineers, lead code reviews, establish engineering standards, and represent platform engineering in cross-functional and executive-level reviews.
Minimum Requirements:
•5 years of hands-on software or platform engineering experience in a production environment, with a builder and developer focus rather than SRE or ticket-driven support.
•4 years of experience developing and configuring the LGTM stack, including Grafana, Mimir, Loki, and Tempo, plus Prometheus for monitoring and alerting.
•5 years of senior-level Python scripting and automation, including exporter development, pipeline scripting, and REST API integrations.
•5 years of GitOps and CI/CD pipeline authoring using GitLab CI/CD, Terraform, and Ansible as primary IaC tools.
•5 years of Linux administration and infrastructure management, including VMware and network fundamentals such as SNMP and TCP/IP.
•2 years of AIOps or observability engineering experience, including Alertmanager rule authoring, anomaly detection, and noise reduction techniques.
•Experience supporting observability across a meaningful infrastructure footprint, targeting at least 1,000-2,000 monitored devices.
Preferred Qualifications
•2 years of secrets management experience using CyberArk, Conjur, HashiCorp Vault, or an equivalent tool with runtime secret injection patterns.
•Experience with ITSM and workflow automation platforms such as ServiceNow, including incident management, alert routing, escalations, and SLA-driven workflows.
•Familiarity with integration platforms such as Boomi or similar iPaaS technologies.
•Hands-on AWS experience building and managing compute, storage, networking, and application services, ideally using multi-account and Well-Architected Framework practices.
•Experience with Building Automation Systems (Client) or Building Management Systems (Client) in data center or OT environments.