Search by job, company or skills

AI Ops Engineer

AI Ops Engineer

D4 Insight
Early Applicant
  • Posted 11 days ago
  • Be among the first 10 applicants

Job Description

Location: Abu Dhabi, UAE

Experience: 7+ Years

Role Overview

We are seeking for AI Ops Engineer to establish the operational backbone for enterprise AI platforms, enabling application teams to release AI products safely, repeatedly, and at scale.

The role is responsible for production release discipline, LLMOps practices, deployment automation, operational governance, cost visibility, and self-service operating standards for AI-native delivery teams.

Key Responsibilities

AI Release & Deployment

  • Build standard release pipelines and promotion controls for AI applications, agents, platform services, and configuration changes.
  • Ensure repeatable deployments with appropriate governance and release evidence.
  • Implement controlled rollout, canary release, and rollback-readiness practices to reduce production risk.

LLMOps & Operational Governance

  • Embed operating controls for models and AI assets to support auditability and AI lifecycle management.
  • Integrate AI quality checks, operational telemetry, dashboards, runbooks, and production-readiness criteria into delivery processes.
  • Establish reusable operational standards and practices for AI-native teams.

Cost & Platform Visibility

  • Provide visibility into AI workload consumption, including:
    • Model usage
    • Token spend
    • Platform capacity
    • Quota management
    • Optimization opportunities
  • Convert proven operating patterns into reusable templates, release standards, onboarding guidance, and operational playbooks.
Required Skills & Experience

  • 7+ years of relevant experience.
  • Strong production engineering background with cloud-native, AI, or high-scale API platforms in enterprise environments.
  • Hands-on experience with CI/CD, GitHub Actions or equivalent automation, and environment management.
  • Working knowledge of:
    • LLMOps
    • Telemetry
    • Release governance
    • Production readiness
  • Experience with observability, change control, service reliability, and continuous operational improvement.
  • Ability to collaborate with Platform Engineering, QA/SRE, Cybersecurity, Architecture, Product, and Delivery teams.
Apply Now: [Confidential Information]

More Info

Key Skills

GitHub Actions

Release governance

Service reliability

High-scale API platforms

LLMOps

CI/CD

Observability

Continuous operational improvement

About Company

Similar Jobs

5-7 yrs
Abu Dhabi, United Arab Emirates
Skills:
AI Release Engineering, GitHub Actions, Change Control, LLMOps, Continuous Operational Improvement, Service Reliability, Observability