Search by job, company or skills

Cloud Operations Engineer

8-10 Years
SGD 6,000 - 9,000 per month
  • Posted 2 days ago
  • Be among the first 10 applicants

Job Description

Role Overview

We are seeking a seasoned Major Incident & Problem Manager to lead high-priority technology incident resolution, post-incident investigations, and operational service continuity across our core banking and infrastructure platforms. You will direct technical command bridges, establish clear operational accountability during crisis events, and ensure rapid service restoration in compliance with regulatory and ITIL standards.

Key Responsibilities

  • Major Incident Command & Control: Take end-to-end ownership of critical incidents (Sev 1 / Sev 2), coordinating across cross-functional engineering, infrastructure, and application teams to minimize Mean Time to Restore (MTTR).
  • Crisis Communication & Governance: Manage executive escalations, provide concise real-time situation updates to senior leadership and business stakeholders, and ensure full compliance with group technology standards.
  • Problem Management & RCA: Facilitate post-incident reviews using structured analysis methodologies (e.g., 5 Whys, Fishbone) to identify underlying causes, eliminate repeat incidents, and track preventative actions to closure.
  • Operational Reporting & Metrics: Monitor incident patterns, compile KPI dashboards (MTTR, SLA compliance, resolution timelines), and support audit and regulatory reporting deliverables.
  • Continuous Service Improvement: Partner with Command Center and Infrastructure units to enhance automated alerting, runbook execution, and incident logging workflows.

Requirements

  • Bachelor's degree in Computer Science or equivalent with around 8 years of relevant experience
  • Proven track record leading Major Incident Management (MIM) and Problem Management within enterprise, high-availability IT environments.
  • Strong working knowledge of ITIL service frameworks (ITIL certification required).
  • Working familiarity with enterprise service management platforms (e.g., BMC Remedy, BMC Helix, ServiceNow).
  • Broad technical literacy across enterprise environments: Application Support, End-of-Day (EOD) batch scheduling, infrastructure components (Linux/Unix, Storage, Network), middleware, and transactional workflows (e.g., payment channels).
  • Exceptional crisis communication and stakeholder management skills, with the ability to maintain composure, command technical calls, and align multiple teams under pressure.
  • Strong analytical capabilities in data reporting and presentation using MS Excel (including macros) and PowerPoint.
  • Flexibility to support critical operational escalations or high-impact release windows when required.

More Info

Job Type:
Industry:
Employment Type:

Job ID: 153715325

Similar Jobs

Singapore

Skills:

GcpTerraformLinuxPowerShellBashWindowsAzurePythonAWS

Singapore

Skills:

GcpTerraformLinuxPowerShellBashWindowsAzurePythonAWS

Singapore

Skills:

RoutingDnsPatch ManagementGcpDevSecOpsTerraformInfrastructure SecurityFirewallsAzureAWSVPCsAccess ControlsDisaster RecoveryAlibaba CloudHybrid ConnectivitypeeringNetwork SegmentationSubnetsVulnerability RemediationInfrastructure-as-CodeBusiness Continuity

Beware of Scammers

We don’t charge money for job offers