Job Summary
Build and optimize the serving layer where scheduling decisions, memory management, and kernel paths determine whether silicon computes or waits. Work at the intersection of machine learning models and hardware acceleration to maximize performance and efficiency.
Responsibilities
- Build and optimize the serving path including batching, memory management, cache behaviour and scheduling under real concurrency.
- Profile end-to-end systems and identify actual bottlenecks rather than assumed ones.
- Enable models to share accelerators effectively across different silicon capacities and generations.
- Work close to the runtime layer including drivers, kernels, memory allocators and telemetry systems.
- Build measurement harnesses and validate results on real hardware.
- Take changes from hypothesis through benchmarking to production deployment.
Must haves
- Four years or more of experience in systems or ML infrastructure engineering.
- Strong proficiency in Python and a systems language such as C++ or Rust.
- Demonstrated experience optimizing inference or training throughput with clear understanding of performance gains.
- Understanding of accelerator memory hierarchies, kernel launch behaviour and performance analysis.
- Comfort working on bare metal rather than behind managed services.
Nice to haves
- Experience with CUDA, ROCm, Triton or comparable kernel-level work.
- Contributions to vLLM, TensorRT-LLM, SGLang or similar projects.
- Experience with distributed serving and multi-accelerator sharding.
- Published benchmark or systems work.
What the company offers
- Position at one of the first deep tech companies in the region building foundational technology in-house.
- Meaningful ownership and impact at an early stage.
- Competitive early-stage compensation.
- Close collaboration with a small, senior team.
- Problems combining hardware, systems and AI at scale.
We refresh listings regularly, but some roles close early on the source platform.
Country: Saudi Arabia
City: Riyadh
Job Category: AI/ML Engineering
Job Type: Full Time
Company Name: think
Seniority level: Mid-Senior level
Sorry! This job has expired.

