Major Incident Manager
الوصف الوظيفي
Major Incident Manager – Drive Operational Resilience and Proactive Service Excellence
As the **Major Incident Manager**, you will play a pivotal role in ensuring operational stability, minimizing business disruption, and driving continuous improvement in incident and problem management. This critical position requires a strategic thinker with exceptional analytical, communication, and leadership skills to lead incident response, root cause analysis, and remediation efforts across complex technical environments. You will be responsible for managing high-impact incidents with urgency, precision, and a focus on long-term service reliability.
In this dynamic role, you will oversee **24x7 incident management**, coordinating cross-functional teams to restore services swiftly while adhering to best practices in ITIL Service Management. Your leadership will be instrumental in establishing clear communication channels, facilitating incident reviews, and driving data-driven improvements to prevent recurring issues. Beyond crisis response, you will proactively identify operational risks, conduct thorough root cause analyses, and implement permanent fixes to enhance service availability and performance.
Key Responsibilities
Your responsibilities will span the entire incident lifecycle, from initial detection to post-mortem analysis and process refinement. Key areas of focus include:
- Incident Leadership and Resolution: Manage critical incidents with urgency, ensuring rapid decision-making and stakeholder alignment to minimize business impact. Lead technical troubleshooting, restoration activities, and communication bridges to maintain transparency during high-pressure situations.
- Incident Review and Documentation: Facilitate **Major Incident Review (MIR) meetings** to analyze incident timelines, triage activities, and recovery actions. Compile detailed reports, documenting lessons learned, root causes, and improvement opportunities to inform future incident response strategies.
- Root Cause Analysis (RCA) and Problem Management: Conduct in-depth RCA for major incidents and recurring issues, collaborating with internal teams, vendors, and OEMs to identify permanent solutions. Proactively register problem records, assess impact trends, and ensure all issues are tracked, resolved, and validated within agreed service levels.
- Proactive Remediation and Service Continuity: Drive initiatives to eliminate recurring problems before they escalate into service-affecting incidents. Implement temporary fixes to restore service continuity while permanent solutions are developed, ensuring compliance with SLAs and KPIs.
- Performance Reporting and Compliance: Produce regular reports on incident trends, RCA status, resolution timelines, and SLA compliance. Monitor KPIs such as Mean Time to Identify Root Cause (MTTRCA), problem closure rates, and proactive issue resolution to measure operational effectiveness.
- Process Improvement and Knowledge Sharing: Identify opportunities to enhance incident and problem management workflows, recommending automation and best-practice enhancements to the Continuous Service Improvement (CSI) team. Update knowledge repositories with verified root causes, workarounds, and permanent resolutions to empower support teams.
- Stakeholder Collaboration: Engage with technical teams, vendors, and stakeholders to ensure alignment on incident response, problem resolution, and process improvements. Foster a culture of accountability and continuous learning to strengthen operational resilience.
Technical and Analytical Expertise
To excel in this role, you will leverage a robust skill set in **Major Incident Management**, **Root Cause Analysis**, and **ITIL Service Management frameworks**. Proficiency in trend analysis, operational analytics, and service performance management will enable you to drive data-informed decisions. Additionally, your technical acumen in infrastructure, applications, cloud services, networks, databases, and security operations will be critical in diagnosing complex issues and implementing effective solutions.
This is an opportunity to make a tangible impact on service reliability, reduce operational risks, and elevate the organization’s incident response capabilities. If you are a detail-oriented leader with a passion for problem-solving and process optimization, we invite you to join our team and help shape the future of our operational excellence.
يمكن أن يرتكب الذكاء الاصطناعي أخطاءً.
ℹ️ إخلاء مسؤولية توظيف:
موقع وظائف السعودية (ksajobshub.com) هو محرك بحث ومجمع لإعلانات الوظائف من المصادر والشركات الرسمية في المملكة العربية السعودية. نحن لا نتقاضى أي مبالغ مالية أو رسوم من الباحثين عن عمل، وتتم عمليات التقديم مباشرة عبر الانتقال للرابط الأصلي للجهة المعلنة.