Job Summary
Develop, deploy, and operate AI/LLM models across dual environments using GCP for public-cloud workloads and Humain sovereign cloud for classified data. Build and optimize models for Arabic NLP, document classification, vision/OCR, and AIOps use cases while ensuring compliance with ZATCA data sovereignty and SDAIA requirements.
Responsibilities
- Build and fine-tune LLM/ML models for Arabic NLP, document classification, vision/OCR, and AIOps applications
- Run pre-deployment evaluation including accuracy baselines, regression testing, and safety testing to justify GPU allocation
- Optimize inference through quantization, batching, and context sizing based on measured usage patterns
- Deploy workloads on Humain GPUaaS using Kubernetes, GPU partitioning on B300 nodes, quotas, and RBAC
- Build equivalent workloads on GCP using Vertex AI and GKE with classification-based routing
- Own serving stack including vLLM/TGI, model versioning, CI/CD, and monitoring for latency, tokens, GPU utilization, and drift
- Ensure developed AI models comply with ZATCA data sovereignty and SDAIA requirements including AI Ethics, GenAI Guidelines, and PDPL
Requirements
- 5 years of ML/AI engineering experience with production LLM deployment
- Proficiency in Python, PyTorch, and Hugging Face
- Production experience with Kubernetes and GPU-served inference
- Experience with GCP Vertex AI or equivalent cloud platform
We refresh listings regularly, but some roles close early on the source platform.
Country: Saudi Arabia
City: Riyadh
Job Category: AI/ML Engineering
Job Type: Full Time
Company Name: InnovationTeam
Seniority level: Mid-Senior level
Sorry! This job has expired.

