Job Summary
As a Site Reliability Engineer, you will ensure high availability and performance of products through proactive support, incident management, and operational automation. You will identify and resolve root causes of operational incidents, manage event catalogs, and develop automated workflows to enhance system reliability and prevent service disruptions.
Responsibilities
- Define, build, and maintain support systems to ensure high availability and performance.
- Implement automation for system provisioning, self-healing, auto-recovery, deployment, and monitoring.
- Perform incident response and root cause analysis (RCA) for critical system failures.
- Monitor system performance and establish Service-Level Indicators (SLIs) and Service-Level Objectives (SLOs).
- Conduct thorough problem investigations and develop permanent solutions for recurring incidents and service disruptions.
- Define, build, and maintain an event catalog specifying active events, thresholds, and remediation actions.
- Collaborate with Development, Operations, and Product teams to integrate reliability best practices and ensure operational readiness for new products.
- Coordinate with Customer Success Managers and internal stakeholders to enhance customer satisfaction and service performance.
Requirements
- Bachelor’s degree in Computer Science, Information Technology, Engineering, or a related field.
- 5+ years of experience in IT operations, service management, or infrastructure management, including roles such as Site Reliability Engineer, Problem Manager, or DevOps Manager.
- Proven experience managing high-availability systems and ensuring operational reliability.
- Extensive experience in root cause analysis (RCA), incident management, and developing permanent solutions for recurring service disruptions.
- Hands-on experience with CI/CD pipelines, automation, system performance monitoring, and infrastructure as code (IaC).
- Strong experience with AKS and on-premises Kubernetes.
- Scripting experience with Ansible, Bash, and Python.
- Experience with Terraform, Azure or AWS, and basic database skills.
- Strong background collaborating with cross-functional teams to improve operational processes and service delivery.
- Familiarity with cloud technologies, containerization, scalable architectures, and zero-downtime deployment strategies.
What the company offers
- Flexible work arrangements: up to 2 days per week from home, flexible daily hours, and up to 30 days per year working from any location globally.
- Employee Assistance Program (EAP) available 24/7, 365 days per year for you and your dependents, plus Champion Health personalized wellbeing platform.
- Professional development through LinkedIn Learning, Microsoft’s Enterprise Skills Initiative, Pluralsight, Harvard Business Publishing, Stanford programs, and Airport Council International resources.
- Competitive benefits aligned with local market and employment status.
We refresh listings regularly, but some roles close early on the source platform.
Country: Jordan
City: Amman
Job Category: Information Technology
Workplace type: Hybrid
Job Type: Full Time
Company Name: SITA
Seniority level: Mid-Senior level
Sorry! This job has expired.

