Job Summary
Evaluate, benchmark, optimize, and validate end-to-end AI inference solutions built on Qualcomm AI accelerators. Analyze real-world customer deployments, assess competitive platforms, and develop insights on performance, scalability, ecosystem readiness, and operational simplicity across modern GenAI workloads including LLMs, VLMs, RAG, and agentic applications.
Responsibilities
- Conduct end-to-end competitive analysis for Qualcomm AI inference accelerators across hardware, software, model-serving frameworks, deployment workflows, and customer-facing GenAI inference solutions.
- Evaluate real customer deployment patterns including LLM/VLM serving, RAG, agentic workflows, embedding services, multi-model applications, and enterprise inference pipelines to identify where Qualcomm leads, is at parity, or has gaps.
- Develop objective benchmark methodologies and execute performance studies covering latency, TTFT, TPOT, throughput, tokens/sec/user, concurrency, utilization, power efficiency, cost efficiency, and single-node and multi-node scaling.
- Compare Qualcomm solutions against competitive platforms such as NVIDIA and AMD across deployment maturity, model coverage, framework support, ecosystem readiness, operational simplicity, and total solution capability.
- Build and maintain reference architectures, deployment playbooks, sizing guidance, reproducible benchmark recipes, and customer-ready solution blueprints for Qualcomm AI inference platforms.
- Partner with architecture, compiler, runtime, framework, systems, and product teams to translate competitive findings into product requirements, roadmap priorities, optimization opportunities, and ecosystem investments.
- Perform system-level debugging and root-cause analysis across model, framework, runtime, driver, hardware, networking, memory, and orchestration layers to explain performance or deployment gaps.
- Create fact-based competitive positioning, leadership readouts, technical reports, and customer-facing collateral that communicate differentiators, risks, gaps, and recommended actions.
Must haves
- Bachelor’s degree in Engineering, Computer Science, or Data Science with 8+ years of experience, OR Master’s degree in Engineering, Computer Science, Data Science or related field with 4+ years of related work experience, OR PhD in Engineering, Computer Science, or related field with 4+ years of experience.
- Strong proficiency in Python and common AI/ML frameworks such as PyTorch, TensorFlow, or ONNX.
- Solid understanding of ML model development, deployment, and AI inference concepts for datacenter workloads.
- Strong foundation in system performance profiling, benchmarking, parallel computing, and workload analysis.
- Working knowledge of Linux-based development, containerized deployment, and cloud-native infrastructure concepts.
- Strong analytical, communication, and cross-functional collaboration skills with the ability to summarize technical findings clearly.
Nice to haves
- Master’s or PhD degree in Engineering, Computer Science, Information Systems, Electrical Engineering, Physics, or a related technical field with 4+ years of experience.
- Strong understanding of GenAI and enterprise inference workloads, including LLMs, VLMs/LVMs, embeddings, diffusion models, RAG pipelines, agents, and multi-turn application patterns.
- Hands-on experience evaluating AI accelerator platforms and deployment ecosystems, including Qualcomm AI inference accelerators and competitive platforms such as NVIDIA and AMD.
- Experience with production inference frameworks and serving stacks such as vLLM, Triton Inference Server, TensorRT-LLM, SGLang, KServe, Ray Serve, TorchServe, or equivalent model-serving technologies.
What the company offers
- Salary including housing and transport allowance, stock (RSUs) and performance related bonus.
- 16 weeks fully paid Maternity Leave and 6 weeks fully paid Paternity Leave.
- Employee stock purchase scheme, Child Education Allowance, and Life and Medical Insurance.
- Relocation and immigration support if needed.
- Live+ Well Reimbursement for health and recreational membership fees.
We refresh listings regularly, but some roles close early on the source platform.
Country: Saudi Arabia
City: Riyadh
Job Category: AI/ML Engineering
Job Type: Full Time
Company Name: Qualcomm
Seniority level: Mid-Senior level
Sorry! This job has expired.

