Search by job, company or skills

DevOps Engineer Platform & AI Infrastructure

DevOps Engineer Platform & AI Infrastructure

Confidential Career Solutions
10-12 Years
Not Disclosed
  • Posted 2 hours ago
  • Be among the first 10 applicants

Job Description

We are building an enterprise AI platform in which AI agents operate business applications through their user interface, the way a person does. It runs on Kubernetes across Microsoft Azure, on-premises, and air-gapped environments, delivered through GitOps. Because it is deployed inside customers own perimeters, it must install and run with no dependency on cloud services in the critical path. We are a small, senior team working in two-week sprints towards a first production release in December 2026.

You will own the target architecture, hosting, and platform: infrastructure as code, cluster management, packaging and delivery, GPU model serving, sandboxed browser workloads, observability, security, and cost. You will lead platform engineering and mentor a junior DevOps engineer.

What you will do

  • Define and own the target architecture and hosting. Produce the architecture and security documentation needed for internal approvals and security clearance.
  • Design and run Kubernetes platforms across Azure (AKS) and on-premises clusters (e.g. RKE2, OpenShift, Rancher, or upstream Kubernetes), with consistent tooling and policies across both.
  • Package the platform as Helm charts that install into customer environments, including air-gapped installs: offline registries (e.g. Harbor), image and chart mirroring, and signed release bundles.
  • Build GitOps delivery with Argo CD or Flux: repository structure, environment promotion, progressive delivery, and drift detection.
  • Define all infrastructure as code with Terraform / OpenTofu and/or Bicep.
  • Run GPU workloads: GPU node pools and scheduling (NVIDIA GPU Operator), and self-hosted serving of open-weight multimodal models (e.g. vLLM, SGLang) behind an OpenAI-compatible gateway.
  • Run fleets of sandboxed headless browsers for AI agents at scale, with strong isolation (e.g. gVisor, Kata Containers, network policies) and controlled egress.
  • Run stateful services on Kubernetes without managed cloud equivalents: PostgreSQL (e.g. CloudNativePG), S3-compatible object storage, and secrets management (e.g. OpenBao / Vault).
  • Own platform security: policy as code (OPA Gatekeeper / Kyverno), image scanning, SBOMs, and supply-chain security.
  • Build observability: metrics, logs, and traces (Prometheus, Grafana, Loki, OpenTelemetry), LLM tracing (e.g. Langfuse), SLOs, and alerting.
  • Own reliability: backup and disaster recovery, capacity planning, incident response, and blameless post-mortems.
  • Lead and mentor the DevOps Engineer – CI/CD & Environments.

What we are looking for

Must have

  • 10+ years in software, infrastructure, DevOps, SRE, or platform engineering, including production Kubernetes at scale.
  • Experience defining target architectures and getting them through security and architecture reviews.
  • Deep hands-on Azure experience (AKS, networking, identity, Key Vault, Monitor, ACR).
  • Experience running on-premises Kubernetes and connecting cloud and on-premises environments.
  • Experience delivering software into air-gapped or highly restricted environments.
  • Production experience with GitOps (Argo CD or Flux), and writing and maintaining Helm charts.
  • Strong infrastructure as code skills with Terraform / OpenTofu and/or Bicep.
  • Solid Linux, networking (TCP/IP, DNS, TLS, load balancing), and scripting (Bash plus Python or Go).
  • A security-first mindset: least privilege, secrets management, policy as code, and supply-chain security.
  • A track record of technical leadership: owning architecture decisions, mentoring engineers, and raising engineering standards across a team.

Nice to have

  • GPU workloads on Kubernetes and LLM serving (vLLM, SGLang, Triton, KServe).
  • Running headless browser fleets (Chromium, Playwright) or other untrusted workloads in sandboxes.
  • Service mesh (Istio, Linkerd, Cilium) and eBPF-based networking.
  • CKA / CKS and/or Azure certifications (AZ-104, AZ-305, AZ-400).
  • Experience in regulated or data-sovereign environments.

What we will assess

  • System design: a platform that runs on Azure and in an air-gapped on-premises site, covering packaging, networking, identity, security, GPU serving, observability, and disaster recovery.
  • Technical exercise: given a service, design its GitOps repository layout, IaC, and promotion flow across environments.
  • Incident scenario: troubleshoot a realistic production failure.
  • Leadership: how you would guide and grow a junior engineer.

Why join

  • Own the architecture of a platform that installs and runs anywhere, from Azure to fully air-gapped sites.
  • Run infrastructure for production AI workloads: GPUs, self-hosted models, and browser sandboxes.
  • Lead platform engineering from day one.

More Info

Job Type:
Industry:
Function:
Employment Type:

Key Skills

Kata Containers

OPA

GitOps

Kyverno

OpenTofu

OpenTelemetry

Bicep

Loki

gVisor

Argo CD