
Search by job, company or skills
Technical Requirements
Experience with infrastructure platforms such as VMware, Hyper-V, or Nutanix
Experience with Red Hat Linux and/or Windows Server environments
Familiarity with enterprise storage, backup, HA, and DR solutions
Strong understanding of monitoring, observability, logging, and alerting practices
Solid understanding of networking fundamentals such as TCP/IP, DNS, routing, firewalls, load balancing, and segmentation
Familiarity with hybrid infrastructure environments, including AWS or Azure
Role Overview
The Platform Operations Engineer supports on-premises and hybrid infrastructure platforms that underpin mission-critical systems. This role focuses on platform reliability, operational stability, and infrastructure improvement in a government or regulated environment.
The engineer will work with an existing internal team and L1 vendor engineers to oversee day-to-day operational support, ensure timely resolution of incidents and issues, and provide deeper technical analysis when required.
The role also includes technical and security governance responsibilities, including reviewing technical changes, supporting service request (SR) review and approval, and ensuring infrastructure changes and operations align with operational, security, and architectural requirements.
This role supports a multi-year technology refresh programme while maintaining stable day-to-day operations with minimal downtime. It also includes technical project coordination responsibilities such as planning, tracking dependencies, managing risks and issues, and driving follow-up across internal teams and vendors to support the agency's longer-term transition towards cloud and hybrid infrastructure.
Key Responsibilities
Work with L1 vendor engineers and the internal team to oversee day-to-day operational support and maintenance of critical infrastructure platforms across on-premises and hybrid environments.
Provide technical oversight across compute, storage, virtualisation, operating systems, backup, disaster recovery (DR), high availability (HA), and related infrastructure platforms.
Lead incident management for complex issues, including technical deep dives, root cause analysis, recovery planning, and timely resolution follow-up.
Manage and coordinate vendor activities to ensure operational tasks, incidents, service requests, and technical deliverables are completed effectively and within expected timelines.
Review and assess technical changes, service requests (SRs), implementation plans, and recovery approaches to ensure they are practical, supportable, and aligned with operational, security, and architectural requirements.
Support technical and security governance activities, including change review, risk awareness, compliance alignment, and operational readiness checks.
Support infrastructure modernisation and the multi-year technology refresh programme while maintaining stable day-to-day operations with minimal downtime.
Plan, coordinate, and track technical infrastructure activities across internal teams and vendors for refresh, upgrade, migration, and other technical project initiatives.
Track project risks, issues, dependencies, action items, and status updates to ensure milestones, operational commitments, and implementation readiness are met.
Work with application, network, security, and vendor teams to resolve platform-related issues and maintain runbooks, SOPs, governance artefacts, and supporting documentation.
Requirements
At least 5 years of experience in infrastructure or platform operations
Experience supporting production or mission-critical environments
Strong technical understanding of infrastructure operations across virtualisation, operating systems, storage, backup, networking, and related platforms
Experience working with vendor-supported operating models and managing external support teams
Able to participate in technical troubleshooting, issue deep dives, and problem resolution for complex incidents
Experience supporting technical governance, security governance, change review, or service request approval processes is preferred
Experience supporting technical project delivery, including coordination, issue and risk tracking, dependency management, and stakeholder follow-up
At least 2 years of experience in government or public sector settings
Good to have data centre operations experience
Strong communication and coordination skills with both technical and non-technical stakeholders
Job ID: 152028177