Job Summary
This role involves supporting application incidents on digital platforms, monitoring and operating the Elastic Observability stack, and collaborating with various teams to ensure system reliability and timely incident resolution.
Responsibilities
- Support application incidents by working with Platform Engineering, Application Development, and customer support teams to meet SLAs and escalation procedures.
- Operate and monitor Elastic Observability components such as Elasticsearch cluster, Kibana, Fleet Server, APM Server, and Elastic Agent managed via ECK on OKE.
- Assist with Elasticsearch operations including index lifecycle management, snapshot lifecycle management, data tier housekeeping, and capacity monitoring.
- Troubleshoot telemetry ingestion issues across logs, metrics, traces, and synthetic monitors to ensure consistent data collection.
- Maintain and update Kibana dashboards, alerting rules, and saved objects under SRE Manager guidance.
- Conduct root cause analysis and participate in blameless post-incident reviews to enhance system reliability.
- Collaborate with Platform Engineering to automate tasks, improve deployment pipelines, and enhance observability using Terraform, Helm charts, and scripting.
- Manage incidents and requests through ticketing systems like Jira and ServiceNow, documenting all activities in the service management system.
Requirements
- Bachelor’s degree in Computer Science, IT, Engineering, or related field, or equivalent experience.
- 1–3 years of experience in IT operations, system administration, application support, DevOps, or SRE.
- Familiarity with Elastic Stack observability tools including Elasticsearch and Kibana.
- Knowledge of Linux systems and scripting languages such as Bash, Python, or Go.
- Understanding of monitoring, logging, and alerting concepts.
- Experience with ITSM tools like ServiceNow, Jira, or Zendesk and ITIL practices.
- Strong knowledge of incident, problem, and change management.
- Basic experience with cloud native environments and containers including Docker and Kubernetes.
- Strong critical thinking, troubleshooting, and communication skills.
We refresh listings regularly, but some roles close early on the source platform.
Country: Saudi Arabia
City: Riyadh
Job Category: Information Technology
Job Type: Full Time
Company Name: Takamol Holding
Seniority level: Mid-Senior level

