Senior Staff Infrastructure & Site Reliability Engineer – Datacentre AI Engineering

Job Summary

Design, deploy, and operate large-scale AI inference systems in a datacenter environment. Ensure reliability, scalability, and production-readiness of Qualcomm’s AI infrastructure while supporting critical machine-learning workloads through strong systems engineering and cross-functional collaboration.

Responsibilities

  • Design, deploy, and operate large-scale AI inference systems supporting critical AI workloads and ensure reliability, availability, and scalability of Qualcomm datacenter AI clusters.
  • Develop and maintain software tools and support infrastructure around AI software stacks.
  • Analyze software requirements and collaborate with architecture and hardware engineers to support AI workloads, including LLM inference and agentic AI workflows.
  • Work with models, systems, and software teams to improve model performance on AI100 deployments and identify optimizations for multi-SoC and multi-card systems.
  • Apply SRE fundamentals including monitoring, alerting, incident response, and performance optimization for production ML systems.
  • Build and maintain observability tools, dashboards, and alerts using Prometheus, Grafana, CloudWatch, and custom telemetry to monitor system health.
  • Develop automation to reduce manual operational tasks and support CI/CD pipelines for AI service deployment using Infrastructure-as-Code tools like Terraform and Ansible.
  • Create and maintain technical documentation, runbooks, and knowledge-base articles for operational guidance.

Must haves

  • 12+ years of software, systems, or infrastructure engineering experience, preferably in production or datacenter environments.
  • Bachelor’s degree in Engineering, Computer Science, AI/ML, or related field (or Master’s degree, or PhD with proportionally fewer years of experience).
  • Strong programming skills in Python with experience building and supporting production systems, plus scripting and automation using Python and Bash.
  • Strong Linux fundamentals including shell, containers, system services, and networking basics (DNS, TLS, HTTP/gRPC).
  • Hands-on experience with monitoring and logging tools such as Prometheus, Grafana, ELK, or Loki.
  • Experience working with AI/ML workloads such as LLMs, NLP, Vision, Audio, or Recommendation systems and hands-on experience with PyTorch.

Nice to haves

  • Experience with GenAI, Agentic AI systems, LLM orchestration frameworks, LangChain, AutoGen, or RAG-based systems.
  • Experience with additional ML frameworks such as TensorFlow, JAX, or Ray.
  • Knowledge of GPU/accelerator-based systems and high-performance networking (RDMA, InfiniBand, RoCE).
  • Experience with advanced MLOps workflows or large-scale AI platform operations.

What the company offers

  • Salary including housing and transport allowance, stock (RSUs), and performance-related bonus.
  • 16 weeks fully paid maternity leave and 6 weeks fully paid paternity leave.
  • Employee stock purchase scheme, child education allowance, and relocation and immigration support if needed.
  • Life and medical insurance, and Live+ Well reimbursement for health and recreational membership fees.

We refresh listings regularly, but some roles close early on the source platform.

Country: Saudi Arabia
City: Riyadh
Job Category: AI/ML Engineering
Job Type: Full Time
Company Name: Qualcomm
Seniority level: Director
Sorry! This job has expired.
Scroll to Top