Purpose:
To design, implement, coordinate, and maintain Disaster Recovery (DR) and Data Center resilience capabilities that ensure the availability, continuity, and rapid restoration of critical IT services, infrastructure, applications, and business operations during disruptive events. The role is responsible for ensuring organizational recovery readiness, compliance with recovery objectives (RTO/RPO), and alignment with business continuity and operational resilience requirements.
Main Duties and Responsibilities:
Disaster Recovery (DR) Management:
- Serve as the Disaster Recovery (DR) Coordinator and primary focal point for all DR-related activities across the organization.
- Develop, maintain, and periodically review Disaster Recovery strategies, policies, plans, procedures, and recovery runbooks.
- Coordinate and manage DR testing exercises, simulation activities, failover and failback tests, and recovery validation programs.
- Ensure recovery objectives (RTO/RPO) are defined, documented, tested, and achieved for critical business services and systems.
- Collaborate with Infrastructure, Application, Security, Business Continuity, and Business Units to maintain DR readiness.
- Conduct DR risk assessments and identify gaps, deficiencies, and improvement opportunities.
- Monitor and track corrective actions resulting from DR tests, audits, assessments, and compliance reviews.
- Maintain DR documentation, inventories, recovery dependencies, and governance records.
- Prepare management reports, dashboards, metrics, and presentations related to DR readiness, testing results, risks, and compliance status.
- Support internal and external audits and ensure compliance with regulatory and organizational requirements.
Data Center Coordination:
- Coordinate day-to-day data center operations to ensure operational stability, service availability, and infrastructure readiness.
- Manage and oversee infrastructure deployment activities, vendor engagements, and operational support requirements.
- Coordinate rack and stack activities, cabling, power allocation, cooling requirements, hardware installation, and equipment relocation activities.
- Ensure adherence to change management processes, operational procedures, and data center governance standards.
- Coordinate infrastructure upgrades, technology refresh projects, migrations, and capacity expansion initiatives.
- Manage operational readiness activities for new systems and infrastructure deployments.
- Support capacity planning, asset management, facility readiness, and operational risk management initiatives.
- Coordinate and monitor third-party vendors to ensure compliance with service delivery and technical requirements.
- Ensure compliance with organizational policies, regulatory standards, and industry best practices related to data center operations.
Project and Operational Resilience Support:
- Lead and coordinate DR and infrastructure resilience projects from planning through implementation.
- Successfully manage strategic initiatives such as:
- Oracle Exadata Cloud Customer (EXACC) Storage Expansion Project.
- Hardware Security Module (HSM) Relocation and Deployment Project.
- Provide technical guidance and recommendations related to recovery capabilities, resilience improvements, and infrastructure optimization.