Job Summary
Design, deploy, and operate large-scale AI inference systems in a datacenter environment. Ensure reliability, scalability, and production-readiness of Qualcomm’s AI infrastructure while supporting critical machine-learning workloads through strong systems engineering and cross-functional collaboration.
Responsibilities
- Design, deploy, and operate large-scale AI inference systems supporting critical AI workloads and ensure reliability, availability, and scalability of Qualcomm datacenter AI clusters.
- Develop and maintain software tools and support infrastructure around AI software stacks.
- Analyze software requirements and collaborate with architecture and hardware engineers to support AI workloads, including LLM inference and agentic AI workflows.
- Work with models, systems, and software teams to improve model performance on AI100 deployments and identify optimizations for multi-SoC and multi-card systems.
- Apply SRE fundamentals including monitoring, alerting, incident response, and performance optimization for production ML systems.
- Build and maintain observability tools, dashboards, and alerts using Prometheus, Grafana, CloudWatch, and custom telemetry to monitor system health.
- Develop automation to reduce manual operational tasks and support CI/CD pipelines for AI service deployment using Infrastructure-as-Code tools like Terraform and Ansible.
- Create and maintain technical documentation, runbooks, and knowledge-base articles for operational guidance.
Must haves
- 12+ years of software, systems, or infrastructure engineering experience, preferably in production or datacenter environments.
- Bachelor’s degree in Engineering, Computer Science, AI/ML, or related field (or Master’s degree, or PhD with proportionally fewer years of experience).
- Strong programming skills in Python with experience building and supporting production systems, plus scripting and automation using Python and Bash.
- Strong Linux fundamentals including shell, containers, system services, and networking basics (DNS, TLS, HTTP/gRPC).
- Hands-on experience with monitoring and logging tools such as Prometheus, Grafana, ELK, or Loki.
- Experience working with AI/ML workloads such as LLMs, NLP, Vision, Audio, or Recommendation systems and hands-on experience with PyTorch.
Nice to haves
- Experience with GenAI, Agentic AI systems, LLM orchestration frameworks, LangChain, AutoGen, or RAG-based systems.
- Experience with additional ML frameworks such as TensorFlow, JAX, or Ray.
- Knowledge of GPU/accelerator-based systems and high-performance networking (RDMA, InfiniBand, RoCE).
- Experience with advanced MLOps workflows or large-scale AI platform operations.
What the company offers
- Salary including housing and transport allowance, stock (RSUs), and performance-related bonus.
- 16 weeks fully paid maternity leave and 6 weeks fully paid paternity leave.
- Employee stock purchase scheme, child education allowance, and relocation and immigration support if needed.
- Life and medical insurance, and Live+ Well reimbursement for health and recreational membership fees.
We refresh listings regularly, but some roles close early on the source platform.
Country: Saudi Arabia
City: Riyadh
Job Category: AI/ML Engineering
Job Type: Full Time
Company Name: Qualcomm
Seniority level: Director
Sorry! This job has expired.

