Role Overview
We are seeking a highly skilled Infrastructure as Code (IaC) Engineer to lead the design, implementation, and management of automated infrastructure provisioning for next-generation AI data centers. This role is central to orchestrating compute, networking, storage, and virtualization layers using Infrastructure as Code (IaC) across on-premises and hybrid cloud environments.
The successful candidate will enable scalable, secure, and repeatable deployment pipelines supporting GPU clusters, AI model training environments, Kubernetes platforms, and high-performance computing (HPC) infrastructure.
Key Responsibilities
Infrastructure Automation
- Design and implement Infrastructure as Code (IaC) frameworks to automate provisioning and lifecycle management of AI data center infrastructure.
- Develop reusable Terraform modules, Ansible playbooks, and YAML templates to standardize infrastructure deployment.
- Automate provisioning and configuration across compute, storage, networking, and virtualization platforms.
- Maintain infrastructure definitions in version-controlled repositories following GitOps best practices.
AI Platform Deployment
- Automate deployment of GPU clusters for AI training and inference workloads.
- Deploy and manage Kubernetes and OpenShift environments.
- Integrate GPU Operators and platform services required for AI/ML environments.
- Support containerized AI platforms and high-performance computing workloads.
Network & Data Center Automation
- Automate provisioning and configuration of:
- Leaf-Spine network architectures
- VXLAN/EVPN fabrics
- BGP routing
- RoCE-enabled AI networking
- Integrate infrastructure automation with physical and virtual networking components.
Storage Automation
- Automate deployment and management of AI storage infrastructure including:
- NVMe storage
- Object Storage (S3-compatible)
- Ceph
- BeeGFS
- Lustre
- Support scalable storage solutions for AI data lakes and high-performance workloads.
CI/CD & DevOps
- Build and maintain Infrastructure CI/CD pipelines using GitLab CI/CD, Jenkins, ArgoCD, or similar tools.
- Implement automated infrastructure testing, validation, and deployment workflows.
- Enable Infrastructure Lifecycle Management through GitOps methodologies.
Monitoring & Observability
- Integrate infrastructure with monitoring and observability platforms including Prometheus, Grafana, and NVIDIA DCGM.
- Implement automated health checks, infrastructure validation, and drift detection.
- Support proactive monitoring and performance optimization of AI infrastructure.
Collaboration & Governance
- Collaborate with AI/ML platform teams, network engineers, security teams, and infrastructure architects.
- Implement infrastructure security standards and policy-as-code frameworks.
- Produce and maintain technical documentation, deployment standards, and operational procedures.
Required Experience
- Bachelor's degree in Computer Science, Information Technology, Engineering, or a related discipline.
- Minimum 5 years of experience in Infrastructure Automation, DevOps, Site Reliability Engineering (SRE), Platform Engineering, or Data Center Infrastructure.
- Strong hands-on experience with Infrastructure as Code (IaC) using Terraform and Ansible.
- Proficiency in scripting languages including Python, Bash, and YAML.
- Experience deploying and managing Kubernetes and/or OpenShift environments.
- Hands-on experience building Infrastructure CI/CD pipelines using GitLab CI/CD, Jenkins, ArgoCD, Flux, or similar platforms.
- Strong understanding of enterprise networking technologies including VXLAN, EVPN, BGP, and RoCE.
- Experience with virtualization platforms such as VMware, KVM, or OpenShift Virtualization.
- Experience using Git and version control best practices.
- Strong troubleshooting, automation, and infrastructure lifecycle management skills.
- Excellent analytical, communication, and collaboration skills.
Preferred Experience
- Experience supporting AI/ML infrastructure or GPU-intensive computing environments.
- Experience deploying NVIDIA GPU infrastructure and GPU Operators.
- Experience automating storage platforms such as Ceph, BeeGFS, Lustre, NVMe, or S3-compatible object storage.
- Experience with monitoring and observability tools including Prometheus, Grafana, NVIDIA DCGM, and Alertmanager.
- Experience implementing GitOps methodologies and Infrastructure Drift Detection.
- Experience with hybrid cloud infrastructure across AWS, Microsoft Azure, or Google Cloud Platform.
- Familiarity with Infrastructure Security, Policy-as-Code, and compliance automation.
- Experience supporting High Performance Computing (HPC) or AI Data Center environments.
Preferred Certifications
- HashiCorp Certified: Terraform Associate
- Red Hat Certified Specialist in Ansible Automation
- Certified Kubernetes Administrator (CKA)
- Red Hat OpenShift Certifications (preferred)
- AWS Certified Solutions Architect, Microsoft Azure Administrator, or Google Cloud Professional certifications (preferred)