Operations Lead
الوصف الوظيفي
Operations Lead - Sakani and Ejar Platforms
The Operations Lead is a critical role responsible for driving the day-to-day operations, service delivery, reliability, and continuous improvement of our Sakani and Ejar platforms. This role ensures platform availability, operational excellence, incident management, disaster recovery readiness, and effective collaboration across various teams and external vendors.
- Operations Management: Lead and manage the operational support of Sakani and Ejar platforms, ensuring platform availability, stability, performance, and reliability. Oversee production and non-production environments, monitor operational KPIs, SLAs, and service health metrics, and drive continuous service improvement initiatives.
- Incident & Problem Management: Lead Major Incident Management activities, coordinate resolution efforts, manage incident escalations, conduct Root Cause Analysis (RCA), and ensure implementation of corrective actions. Track recurring issues and drive long-term solutions.
- Release & Change Management: Oversee production deployments, releases, and maintenance activities. Ensure operational readiness before major releases, review implementation and rollback plans, and ensure compliance with change management processes and governance standards.
- Platform & DevOps Operations: Collaborate with DevOps teams to enhance automation, reliability, and operational efficiency. Ensure monitoring, logging, alerting, and observability capabilities are maintained, and support Kubernetes, cloud infrastructure, CI/CD pipelines, and platform operations. Drive operational excellence through automation and process optimization.
- Disaster Recovery & Business Continuity: Lead Disaster Recovery (DR) planning, testing, and execution activities. Ensure Recovery Time Objective (RTO) and Recovery Point Objective (RPO) targets are achieved, maintain operational runbooks and recovery procedures, and coordinate periodic DR drills and readiness assessments.
- Stakeholder Management: Act as the primary operational contact for business and technical stakeholders. Coordinate with internal teams, external vendors, and service providers, prepare and present operational reports, service reviews, and executive updates, and facilitate operational governance and service review meetings.
- Vendor & Service Management: Manage vendor performance against agreed Service Level Agreements (SLAs), coordinate operational activities with third-party providers, and handle escalations to ensure timely issue resolution.
- Performance & Capacity Management: Monitor platform performance, utilization, and capacity trends, identify and address performance bottlenecks, and plan future capacity requirements and scalability improvements.
- Security, Risk & Compliance: Ensure compliance with organizational policies, security standards, and regulatory requirements. Support security audits, risk assessments, and compliance initiatives, and track and mitigate operational risks.
Key Performance Indicators (KPIs): Platform Availability ≥ 99.9%, SLA Compliance ≥ 95%, Change Success Rate ≥ 98%, Reduction in Mean Time to Recovery (MTTR), Successful Disaster Recovery testing and execution, Operational Risk Reduction, Stakeholder and Customer Satisfaction.
Reporting To: IT Operations Manager / Head of Technology Operations.
Scope: Responsible for operational governance, service reliability, incident management, platform operations, disaster recovery readiness, and service delivery across.
Qualifications: Bachelor’s degree in Computer Science, Information Technology, Engineering, or a related field, with a minimum of 8 years of experience in IT Operations, Service Delivery, Infrastructure, or DevOps, and at least 3 years of experience in a leadership or management role. Proven experience managing large-scale enterprise platforms and critical business services, working with cross-functional teams, and external vendors.
يمكن أن يرتكب الذكاء الاصطناعي أخطاءً.