Sr Platform Engineer
About this role
Employer-provided description, formatted for easier reading.
Sr. Platform EngineerSummary of Position
Your goal: deploy our cloud applications to production several times a day, with automated tests gating every release, automatic failback when a release goes wrong, alerts that fire on their own, systems that heal themselves, and disaster recovery we have tested.
You’ll do the hands-on work and own everything between a commit and running software: GitLab CI/CD, infrastructure as code on AWS, the QA and demo environments, and the monitoring and paging that catch a failure before a customer does. Our older hosted products have their own operations team, so you won’t support client installs.
Using AI well is part of the job; you’ll build with Claude Code and AI agents writing and reviewing much of the pipeline, template, and automation code while you own the result.
This role will work closely across engineering, IT, security, and other stakeholders, so strong communication skills and the ability to collaborate effectively across teams are essential. As the platform function grows, there may also be an opportunity to take on additional technical leadership responsibilities in the future.
Role and Responsibilities
- Get to several deploys a day:
Build the path from commit to production, with automated tests as the gate and automatic failback when a release goes bad.
- Work AI-first:
Use Claude Code, agent skills, and MCP servers to write, test, and review pipelines, templates, and automation.
- Own CI/CD:
GitLab pipelines, build reliability, and automated deploys into every AWS environment.
- Own infrastructure as code:
Bring every AWS account under version control, CloudFormation today, so nothing changes by console click.
- Run the AWS estate jointly with IT:
Account structure, IAM, networking, and cost; find waste, right-size, and report spend against budget every month.
- Build the environments engineering works in:
QA, demo, and UAT, and stand up a fresh one on request instead of by hand.
- Make the systems heal themselves:
Health checks, auto-scaling, and automatic restarts, so most failures recover before anyone is paged.
- Own monitoring and paging:
Wire metrics, logs, traces, and alarms so an incident shows up in monitoring before it shows up as a support call; support the PagerDuty rotation and help ensure alerts are actionable.
- Own disaster recovery:
Run quarterly recovery tests against a written restore-time target and ticket every gap.
- Carry infrastructure’s side of security and compliance:
SOC 2 evidence, pen-test fixes, secrets management, encryption at rest, and patching; the Director of Security owns the overall program.
- Collaborate across teams:
Communicate clearly with engineering, IT, security, and other stakeholders to troubleshoot issues, drive infrastructure initiatives, and keep work moving.
- Stay hands-on:
Build the pipelines, templates, and automation yourself.
Qualifications and Education Requirements
- Bachelor’s Degree in Computer Science or related field.
- 7+ years in DevOps, SRE, or platform engineering, with experience owning a cloud estate in production.
- Experience building or supporting continuous delivery: several production deploys a day, automated test gates, and automatic failback.
- Daily use of AI coding agents such as Claude Code or Cursor on infrastructure work; able to explain what you shipped faster with them, where the agent went wrong, and how you caught it.
- AWS in depth: Fargate and ECS, Aurora, S3, CloudFront, API Gateway, Cognito, SQS and SNS, Lambda, WAF, IAM, and CloudWatch.
- Infrastructure as code in production: CloudFormation or Terraform, with the whole estate managed there rather than partially by hand.
- Experience owning CI/CD at scale, GitLab CI preferred: pipeline design, build reliability, and automated deploys across several environments.
- Docker in production, orchestrated on Fargate, ECS, or Kubernetes.
- Experience standing up metrics, alerting, and on-call from nothing and tuning it so the pages that fire are meaningful.
- Scripts in Python, Bash, or PowerShell, and reads Java and Angular build tooling well enough to debug a pipeline.
- Strong communication and collaboration skills, with the ability to work effectively across engineering, IT, security, and other technical and non-technical stakeholders.
- Treats security as part of the job: secrets management, least-privilege IAM, regular patching, and evidence an auditor accepts.
- Nice to have:
PagerDuty; Aurora MySQL operations; database migrations and encryption-at-rest rollouts across many customer databases; building MCP servers or agent skills.