Associate Staff Engineer(SRE)
What you'll need to apply
What this employer's standard application typically asks
About this role
Employer-provided description, formatted for easier reading.
Must have Skills : DevOps - AWS (Strong)
Job Description :
Senior Site Reliability Engineer (SRE) Role
We are seeking an experienced Senior Site Reliability Engineer (SRE) to support highly available, business-critical platforms within a global hospitality environment.
The role is responsible for ensuring the reliability, availability, performance, scalability, security, and operational readiness of production systems supporting hotel operations, reservation services, digital channels, loyalty platforms, payment services, and other guest-facing applications.
The successful candidate will combine strong expertise in Cloud, DevOps, Kubernetes, Infrastructure as Code, Observability, Automation, and Production Support, and will work closely with Development, Infrastructure, Security, Architecture, and Service Management teams in a global 24x7 operating model.
Key Responsibilities
- Ensure the availability, reliability, performance, and scalability of critical production services.
- Define and monitor SLIs, SLOs, SLAs, error budgets, availability, latency, and service health metrics.
- Act as a senior technical escalation point for major production incidents and P1/P2 issues.
- Lead troubleshooting, Root Cause Analysis, post-incident reviews, and corrective actions.
- Reduce operational toil through automation, self-healing, and engineering improvements.
- Operate and troubleshoot workloads across AWS, Azure, and/or GCP environments.
- Support enterprise Kubernetes and container platforms, including EKS, AKS, GKE, or OpenShift.
- Develop and maintain Infrastructure as Code using Terraform, CloudFormation, Bicep, or equivalent technologies.
- Build and maintain CI/CD pipelines using tools such as Jenkins, GitHub Actions, GitLab CI, Azure DevOps, Argo CD, or Harness.
- Implement and maintain observability solutions covering metrics, logs, tracing, alerts, dashboards, and APM.
- Support tools such as Prometheus, Grafana, Dynatrace, Datadog, Splunk, ELK/OpenSearch, or New Relic.
- Perform performance analysis, capacity planning, load testing, and scalability optimization.
- Support Disaster Recovery, Business Continuity, failover testing, and RTO/RPO validation.
- Participate in production release, change, patching, vulnerability remediation, and operational readiness activities.
- Maintain technical runbooks, operational procedures, monitoring standards, and knowledge documentation.
- Participate in a global 24x7 on-call / production support model where required.
Hospitality Technology Scope The role may support platforms including
Property Management Systems (PMS) Central Reservation Systems (CRS) Hotel booking and reservation platforms Loyalty and membership systems Guest-facing websites and mobile applications Payment platforms Property connectivity and integration services API and middleware platforms Revenue management and hotel operational systems
The engineer will help ensure the reliability of critical guest journeys such as hotel search, booking, reservation modification, payment, check-in/check-out, loyalty transactions, and property system integration.
Required Qualifications
- Bachelor's degree in Computer Science, Engineering, Information Technology, or a related discipline.
- 7+ years of experience in Cloud, DevOps, Infrastructure, Production Engineering, or IT Operations.
- 3+ years of hands-on experience in an SRE, DevOps, Cloud Operations, or Production Engineering role.
Strong hands-on experience with at least one major cloud platform
- AWS, Azure, or GCP.
- Strong Kubernetes, Docker, and container troubleshooting skills.
- Hands-on experience with Terraform or other Infrastructure as Code technologies.
- Strong experience with CI/CD, deployment automation, and release management.
- Strong knowledge of Linux and production troubleshooting.
- Experience with enterprise monitoring, logging, tracing, and observability platforms.
- Experience managing major incidents, RCA, problem management, and production stability.
- Good knowledge of networking concepts including DNS, HTTP/HTTPS, load balancing, firewall, routing, VPN, and CDN.
- Experience supporting microservices, distributed systems, APIs, databases, and messaging platforms.
- Scripting or programming experience with Python, Bash, PowerShell, Go, Java, or similar languages.
- Good understanding of ITIL-based Incident, Problem, Change, and Knowledge Management processes.
- Strong written and verbal English communication skills.
Preferred Qualifications
- Experience in hospitality, travel, airline, e-commerce, financial services, or other 24x7 high-availability industries.
- Experience supporting high-volume transactional or reservation platforms.
- Experience with PCI DSS, GDPR, ISO 27001, DevSecOps, and vulnerability management.
- Experience with Kafka, Redis, API gateways, service mesh, or event-driven architectures.
- Experience with cloud cost optimization / FinOps.
- Experience with resilience testing, chaos engineering, or automated recovery.
- Relevant certifications such as AWS, Azure, GCP, CKA, Terraform Associate, or ITIL.