Key Responsibilities
- Perform 24x7 incident management activities to minimize service disruption and business impact.
- Register and manage incidents resulting from service degradation, outages, or operational alerts.
- Establish and manage communication, Incident Bridges for critical incidents.
- Coordinate technical troubleshooting, restoration activities, and stakeholder engagement.
- Ensure rapid decision-making and removal of operational roadblocks during incident resolution.
- Drive service restoration efforts while maintaining adherence to Incident Management processes
- Schedule and facilitate Major Incident Review meetings following incident resolution.
- Document incident timelines, triage activities, recovery actions, lessons learned, and improvement opportunities.
- Coordinate stakeholder participation during review sessions.
- Finalize and publish MIR reports.
- Coordinate Root Cause Analysis (RCA) activities for major incidents and recurring operational issues across internal technical teams, OEMs, third-party vendors, and stakeholders.
- Proactively identify and register problem records based on incident trends and potential service risks.
- Perform detailed problem determination and impact assessment activities.
- Create and manage problem records for all major incidents and significant service-impacting events.
- Ensure all problem records contain documented root causes, workarounds, corrective actions, and permanent resolutions.
- Track open problems and monitor progress toward resolution within agreed service levels.
- Facilitate investigation activities to determine underlying causes of incidents.
- Coordinate with technical teams, vendors, OEMs, and SAMA stakeholders to implement permanent fixes.
- Drive proactive remediation initiatives to eliminate recurring issues before they become service-affecting trends.
- Ensure corrective and preventive actions are implemented and validated.
- Coordinate implementation of temporary fixes to restore service continuity while permanent solutions are developed.
- Validate effectiveness of workarounds and permanent resolutions.
- Produce regular reports as agreed KPI like problem trends, RCA status, resolution timelines, reoccurring problems, incidents, SLA compliance etc.
- Ensure compliance with Incident and Problem Management policies, operational procedures, and regulatory requirements.
- Identify opportunities to enhance the Incident and Problem Management process.
- Provide recommendations to the CSI team regarding process improvements and automation opportunities.
- Drive initiatives aimed at reducing incident recurrence and improving service availability.
- Update knowledge repositories with verified root causes, known errors, workaround procedures, and permanent resolutions.
- Ensure lessons learned from major incidents are documented and communicated across support teams.
Technical Skills
- Major Incident Management
- Strong Root Cause Analysis (RCA) methodologies.
- Incident and Problem Management practices.
- Knowledge of ITIL Service Management framework.
- Trend analysis, reporting, and operational analytics.
- Knowledge of infrastructure, applications, cloud services, networks, databases, and security operations.
- SLAs, KPIs, and service performance management.
KPIs
- Reduction in recurring incidents.
- Percentage of Major Incidents with complete RCA.
- Problem and Major incident closure within SLA targets.
- Mean Time to Identify Root Cause (MTTRCA).
- Number of proactive problems identified and resolved.
Reduction in incident volume through permanent fixes