Site Reliability Engineer Manager (SREM)

Job Summary

We are seeking a Senior Site Reliability Engineer to lead reliability and performance initiatives across the engineering organization. You will apply software engineering principles to system administration, building fault-tolerant infrastructure, establishing service level objectives, and ensuring platform availability and security while meeting KSA financial and data compliance requirements.

Responsibilities

  • Define, measure, and manage Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Error Budgets to balance feature velocity with system stability
  • Design, build, and maintain cloud-native infrastructure on AWS, GCP, or Azure using Infrastructure as Code tools such as Terraform or Pulumi
  • Build automation tools to eliminate manual work in provisioning, scaling, failover, and deployment processes
  • Lead incident management processes, participate in on-call rotations, and conduct blameless post-mortems to prevent recurring issues
  • Implement comprehensive monitoring, logging, and alerting systems using tools like Prometheus, Grafana, Datadog, or ELK stack
  • Optimize deployment pipelines with development teams for speed, safety, and reliability
  • Collaborate with security teams to implement security practices and ensure infrastructure compliance with KSA regulations including SAMA, NCA, and PDPL data localization requirements
  • Mentor junior engineers and foster a culture of reliability and operational excellence

Must haves

  • 5+ years of experience in Software Engineering, DevOps, or SRE roles, with at least 2 years in a senior capacity
  • Deep, hands-on experience with major cloud providers, preferably AWS or GCP
  • Strong proficiency in container orchestration, specifically Kubernetes and Docker
  • Strong programming skills in at least one language such as Go, Python, or Java
  • Proven experience with Infrastructure as Code tools such as Terraform or Ansible
  • Deep understanding of Linux operating systems, networking (TCP/IP, DNS, routing), and distributed systems architecture
  • Excellent verbal and written communication skills in English, with Arabic as a strong plus

Nice to haves

  • Prior experience working in a Fintech, Proptech, or highly regulated environment
  • Familiarity with Saudi Arabian financial and data privacy regulations (SAMA, NCA frameworks)
  • Experience with database reliability including PostgreSQL, Redis, and Kafka

We refresh listings regularly, but some roles close early on the source platform.

Country: Saudi Arabia
City: Riyadh
Job Category: Software Engineering
Job Type: Full Time
Company Name: Soar
Seniority level: Mid-Senior level
Sorry! This job has expired.
Scroll to Top