Search by job, company or skills

infrastructure as Code

  • Posted 9 hours ago
  • Be among the first 10 applicants

Job Description

Role Overview

We are seeking a highly skilled Infrastructure as Code (IaC) Engineer to lead the design, implementation, and management of automated infrastructure provisioning for next-generation AI data centers. This role is central to orchestrating compute, networking, storage, and virtualization layers using Infrastructure as Code (IaC) across on-premises and hybrid cloud environments.

The successful candidate will enable scalable, secure, and repeatable deployment pipelines supporting GPU clusters, AI model training environments, Kubernetes platforms, and high-performance computing (HPC) infrastructure.

Key Responsibilities

Infrastructure Automation

  • Design and implement Infrastructure as Code (IaC) frameworks to automate provisioning and lifecycle management of AI data center infrastructure.
  • Develop reusable Terraform modules, Ansible playbooks, and YAML templates to standardize infrastructure deployment.
  • Automate provisioning and configuration across compute, storage, networking, and virtualization platforms.
  • Maintain infrastructure definitions in version-controlled repositories following GitOps best practices.

AI Platform Deployment

  • Automate deployment of GPU clusters for AI training and inference workloads.
  • Deploy and manage Kubernetes and OpenShift environments.
  • Integrate GPU Operators and platform services required for AI/ML environments.
  • Support containerized AI platforms and high-performance computing workloads.

Network & Data Center Automation

  • Automate provisioning and configuration of:
  • Leaf-Spine network architectures
  • VXLAN/EVPN fabrics
  • BGP routing
  • RoCE-enabled AI networking
  • Integrate infrastructure automation with physical and virtual networking components.

Storage Automation

  • Automate deployment and management of AI storage infrastructure including:
  • NVMe storage
  • Object Storage (S3-compatible)
  • Ceph
  • BeeGFS
  • Lustre
  • Support scalable storage solutions for AI data lakes and high-performance workloads.

CI/CD & DevOps

  • Build and maintain Infrastructure CI/CD pipelines using GitLab CI/CD, Jenkins, ArgoCD, or similar tools.
  • Implement automated infrastructure testing, validation, and deployment workflows.
  • Enable Infrastructure Lifecycle Management through GitOps methodologies.

Monitoring & Observability

  • Integrate infrastructure with monitoring and observability platforms including Prometheus, Grafana, and NVIDIA DCGM.
  • Implement automated health checks, infrastructure validation, and drift detection.
  • Support proactive monitoring and performance optimization of AI infrastructure.

Collaboration & Governance

  • Collaborate with AI/ML platform teams, network engineers, security teams, and infrastructure architects.
  • Implement infrastructure security standards and policy-as-code frameworks.
  • Produce and maintain technical documentation, deployment standards, and operational procedures.

Required Experience

  • Bachelor's degree in Computer Science, Information Technology, Engineering, or a related discipline.
  • Minimum 5 years of experience in Infrastructure Automation, DevOps, Site Reliability Engineering (SRE), Platform Engineering, or Data Center Infrastructure.
  • Strong hands-on experience with Infrastructure as Code (IaC) using Terraform and Ansible.
  • Proficiency in scripting languages including Python, Bash, and YAML.
  • Experience deploying and managing Kubernetes and/or OpenShift environments.
  • Hands-on experience building Infrastructure CI/CD pipelines using GitLab CI/CD, Jenkins, ArgoCD, Flux, or similar platforms.
  • Strong understanding of enterprise networking technologies including VXLAN, EVPN, BGP, and RoCE.
  • Experience with virtualization platforms such as VMware, KVM, or OpenShift Virtualization.
  • Experience using Git and version control best practices.
  • Strong troubleshooting, automation, and infrastructure lifecycle management skills.
  • Excellent analytical, communication, and collaboration skills.

Preferred Experience

  • Experience supporting AI/ML infrastructure or GPU-intensive computing environments.
  • Experience deploying NVIDIA GPU infrastructure and GPU Operators.
  • Experience automating storage platforms such as Ceph, BeeGFS, Lustre, NVMe, or S3-compatible object storage.
  • Experience with monitoring and observability tools including Prometheus, Grafana, NVIDIA DCGM, and Alertmanager.
  • Experience implementing GitOps methodologies and Infrastructure Drift Detection.
  • Experience with hybrid cloud infrastructure across AWS, Microsoft Azure, or Google Cloud Platform.
  • Familiarity with Infrastructure Security, Policy-as-Code, and compliance automation.
  • Experience supporting High Performance Computing (HPC) or AI Data Center environments.

Preferred Certifications

  • HashiCorp Certified: Terraform Associate
  • Red Hat Certified Specialist in Ansible Automation
  • Certified Kubernetes Administrator (CKA)
  • Red Hat OpenShift Certifications (preferred)
  • AWS Certified Solutions Architect, Microsoft Azure Administrator, or Google Cloud Professional certifications (preferred)

More Info

Job Type:
Industry:
Employment Type:

Job ID: 151784275

Beware of Scammers

We don’t charge money for job offers