
Search by job, company or skills
Position: Senior AI Engineer
Compensation: Competitive salary plus family benefits & variable
Location: Dubai, United Arab Emirates
Overview
Gateworth Group is partnering with a technology organisation in the UAE that is scaling its AI engineering function and building a specialist team focused on high‑performance model optimisation. They're investing heavily in advanced AI infrastructure and are seeking senior engineers who can work deep inside modern LLMs, improve inference behaviour, and shape how large‑scale models run in production.
This role suits someone who enjoys complex, hands‑on engineering, understands how transformer architectures behave at scale, and can move confidently between model internals, systems performance, and deployment‑level optimisation. You'll work across analysis, tuning, benchmarking and architectural decision‑making, contributing to next‑generation AI systems.
Main Responsibilities
• Improve and optimise LLM inference performance across distributed, multi‑chip and multi‑node environments
• Apply strong understanding of transformer architectures, including dense and Mixture‑of‑Experts (MoE) models
• Benchmark leading LLMs (LLaMA, Mistral, Qwen, DeepSeek) across varied hardware stacks
• Design and implement attention‑level optimisations (Flash Attention, grouped‑query, sliding‑window)
• Deliver model‑level optimisation including quantisation (INT8/FP8), KV‑cache strategies, batching and parallelism
• Work closely with hardware, systems and compiler teams to co‑design efficient inference pipelines
• Build and maintain benchmarking frameworks to measure latency, throughput and scaling behaviour
• Evaluate architectural trade‑offs and contribute to deployment strategies for large‑scale environments
• Stay current with research across LLM architectures, inference optimisation and performance engineering
Qualifications
• Strong understanding of transformer architectures, LLM internals and both dense/MoE models
• Hands‑on experience with modern LLMs (LLaMA, Mistral, Qwen, DeepSeek) and attention‑level optimisation
• Practical experience with inference optimisation: quantisation (INT8/FP8), KV‑cache strategies, batching, pruning and parallelism
• Strong Python skills with PyTorch or JAX, plus experience profiling and debugging performance bottlenecks
• Background in distributed systems, large‑scale inference workloads and system‑level optimisation across hardware and runtime layers
• Ideally 8+ years in deep learning, AI systems or performance engineering, with exposure to datacenter‑scale inference (e.g., vLLM) and hardware‑aware optimisation
Apply
Candidates who meet the above criteria and are seeking a progressive, technically challenging role are invited to apply via the link provided. For further questions, contact [Confidential Information] quoting reference #8341.
Job ID: 153599071
Skills:
AI ML, Python, Llm, RAG, Gen AI, AI Engineer
Skills:
Bash, Cdn, Gcp, Terraform, Siem, Waf, Azure, Python, AWS, XDR, AI-driven anomaly detection, EDR, UEBA
Skills:
Automated Testing, Containers, Typescript, React, FastAPI, Python, Apis, Git, Azure ML, Next.js, tool calling, Search, GenAI, Azure AI Foundry, Azure OpenAI, vector databases, embeddings, CI CD, agent orchestration, Azure AI Search, LLM applications, RAG AI agents, prompt engineering, enterprise integrations
Skills:
Backend Engineering, Python, Apis, Typescript, React, MCP servers, modern frontend tooling, AI evaluation tracing tools, orchestration patterns, modern LLM frameworks, agentic AI applications, external system integrations
Skills:
ECS, Ml, Golang, MLops, AWS, Node.js, Python, Docker, Jenkins, GitHub Actions, EKS, Ai, vector databases, NoSQL databases