Search by job, company or skills

IT Administrator HPC & Data Center Operations

Early Applicant
  • Posted 9 hours ago
  • Be among the first 10 applicants

Job Description

Job Description: IT Administrator — HPC & Data Center Operations

Department: Information Technology / Infrastructure

Reports To: Head of IT / Chief Technology Officer

Location: Abu Dhabi - On-site (Data Center presence required)

Employment Type: Full-Time

About the Role

Prepaire Labs is seeking an experienced IT Administrator to manage and

maintain our High-Performance Computing (HPC) environment, including CPU

and GPU data center infrastructure that powers our AI-driven drug discovery and

healthcare research platforms. The successful candidate will be responsible for

end-to-end administration of compute clusters, networking connectivity,

healthcare and scientific software licensing, and day-to-day IT operations,

ensuring high availability, security, and regulatory compliance across all systems

Key Responsibilities

1. HPC & Data Center Administration (CPU/GPU)

• Administer, monitor, and maintain HPC clusters comprising CPU and GPU

compute nodes (e.g., NVIDIA H100/A100/L40S, AMD EPYC, Intel Xeon

platforms).

• Deploy, configure, and manage cluster workload managers and job

schedulers (SLURM, PBS, or Kubernetes with GPU orchestration).

• Manage GPU drivers, CUDA toolkits, container runtimes (Docker,

Singularity/Apptainer, NVIDIA Container Toolkit), and ML frameworks

supporting research workloads.

• Oversee data center operations: rack and stack, power and cooling

monitoring (PDU/UPS), cable management, hardware lifecycle, and capacity

planning.

• Perform preventive maintenance, firmware/BIOS updates, hardware

diagnostics, and coordinate vendor RMA/support cases (NVIDIA, Dell, HPE,

Supermicro, Lenovo, etc.).

• Manage high-performance storage systems (NFS, Lustre, BeeGFS, Ceph, or

enterprise NAS/SAN) and data backup/disaster recovery strategies.

• Monitor cluster health, utilization, and performance using tools such as

Grafana, Prometheus, Zabbix, Nagios, or NVIDIA DCGM; optimize resource

allocation for research teams.

2. Networking & Connectivity

Design, configure, and maintain LAN/WAN, VLANs, firewalls, VPNs, and

wireless infrastructure across office and data center environments.

• Administer high-speed interconnects for HPC workloads (InfiniBand, RoCE,

10/25/40/100GbE).

• Manage core network equipment (Cisco, Juniper, Arista, Fortinet, Mellanox/

NVIDIA networking) including switching, routing, and firmware updates.

• Ensure secure remote access, site-to-site connectivity, and redundancy/

failover for critical links.

• Monitor network performance, troubleshoot latency/bandwidth issues, and

maintain network documentation, IP address management (IPAM), and

topology diagrams.

• Implement network segmentation and access controls appropriate for

sensitive healthcare and research data.

3. Healthcare Network, Software & Licensing Management

• Manage procurement, deployment, renewal, and compliance of software

licenses, including scientific, healthcare, and laboratory applications (e.g.,

LIMS, bioinformatics suites, molecular modeling tools, EHR/EMR

integrations where applicable).

• Administer license servers (FlexLM/FlexNet, RLM, etc.) for engineering and

scientific software.

• Track license utilization, forecast needs, and optimize licensing costs across

teams.

• Ensure all systems handling healthcare or patient-related data comply with

applicable regulations and standards (HIPAA, GDPR, ISO 27001, GxP/21 CFR

Part 11 as relevant).

• Liaise with healthcare technology vendors, managed service providers, and

regulatory/compliance teams.

4. Systems & Security Administration

• Administer Linux (RHEL/Rocky/Ubuntu) and Windows Server

environments, including Active Directory / LDAP, DNS, DHCP, and identity/

access management (SSO, MFA).

• Implement and enforce security policies: patch management, endpoint

protection, vulnerability scanning, intrusion detection, and audit logging.

• Manage virtualization and cloud resources (VMware, Proxmox, Hyper-V;

AWS/Azure/GCP hybrid connectivity where required).

• Maintain robust backup, replication, and disaster recovery procedures;

conduct periodic DR testing.

• Support incident response, root cause analysis, and change management processes.

5. IT Operations & User Support

• Provide Tier 2/3 support for researchers, scientists, and staff on workstations,

peripherals, collaboration tools, and access to HPC resources.

• Onboard/offboard users, manage accounts, permissions, and research

project allocations on compute clusters.

• Maintain accurate documentation: SOPs, runbooks, asset inventory, and

configuration records.

• Train end users on responsible and efficient use of HPC and IT resources.

• Participate in on-call rotation for critical infrastructure issues.

Required Qualifications

• Bachelor's degree in Computer Science, Information Technology,

Engineering, or equivalent practical experience.

• 5+ years of experience in IT systems administration, with 2+ years

managing HPC and/or GPU data center environments.

• Strong hands-on expertise in Linux administration (RHEL/CentOS/Rocky/

Ubuntu) and shell scripting (Bash, Python).

• Proven experience with GPU computing stacks (NVIDIA drivers, CUDA,

DCGM) and job schedulers (SLURM preferred).

• Solid networking knowledge: TCP/IP, VLANs, routing, firewalls, VPNs, and

high-speed interconnects (InfiniBand/100GbE a plus).

• Experience managing software licensing, license servers, and vendor

contracts.

• Familiarity with healthcare data compliance requirements (HIPAA, GDPR)

and IT security best practices.

• Strong troubleshooting, documentation, and communication skills.

Preferred Qualifications

• Certifications such as RHCSA/RHCE, CCNA/CCNP, CompTIA Server+/

Security+, NVIDIA-certified (e.g., NCA/NCP for AI Infrastructure), VMware

VCP, or ITIL Foundation.

• Experience supporting biotech, pharma, healthcare, or research laboratory

environments (LIMS, bioinformatics pipelines).

• Experience with infrastructure-as-code and automation (Ansible,

Terraform, Git).

• Exposure to container orchestration (Kubernetes, Helm) and MLOps

environments.

• Knowledge of data center standards (Tier classifications, hot/cold aisle

design, DCIM tools).

• Experience with hybrid cloud HPC bursting and cost optimization.

Key Competencies

• Ownership mindset with the ability to manage critical 24/7 infrastructure.

• Analytical problem-solver who thrives in fast-paced R&D environments.

• Detail-oriented approach to compliance, documentation, and change

control.

• Collaborative communicator able to translate technical matters for scientific

and business stakeholders.

• Proactive about capacity planning, performance tuning, and continuous

improvement.

What We Offer

• Opportunity to work at the intersection of AI, HPC, and healthcare

innovation.

• Competitive salary and benefits package.

• Professional development, training, and certification support.

• A collaborative, mission-driven research environment.

More Info

Job Type:
Industry:
Employment Type:

About Company

Job ID: 151874207

Beware of Scammers

We don’t charge money for job offers