Digital Infrastructure & Cloud Lead
الوصف الوظيفي
Role Purpose
As the Digital Infrastructure & Cloud Lead, you will be responsible for designing, deploying, and maintaining the organization’s cloud and on-premises infrastructure platforms. Your role is pivotal in ensuring these environments are stable, secure, and fully reproducible through hands-on provisioning, Infrastructure-as-Code (IaC), automation, and observability practices. You will empower the infrastructure team by reducing manual effort, resolving complex incidents, and fostering resilient operations that enable seamless application delivery. Your expertise will drive efficiency, reliability, and scalability across the digital infrastructure landscape.
Key Responsibilities
Infrastructure & Cloud Provisioning and Deployment
You will implement and maintain cloud and on-premises infrastructure using standardized patterns and IaC templates, ensuring the provisioning of resources, networks, storage, and platform services. Your responsibilities include deploying and validating environments across development, testing, staging, and production stages with repeatable pipelines. By maintaining environment parity and reproducible builds, you will guarantee consistency and reliability throughout the infrastructure lifecycle.
Automation, Configuration & Infrastructure as Code
You will develop and maintain IaC modules, automation scripts, and deployment pipelines to minimize manual interventions, enhance repeatability, and support self-service provisioning capabilities. Additionally, you will create and maintain automated operational playbooks for tasks such as configuration drift remediation, automated patching, and self-healing processes. These initiatives will significantly reduce operational toil and accelerate incident recovery times.
Monitoring, Observability & Performance Tuning
You will implement comprehensive observability solutions, including metrics, logs, and traces, while building intuitive dashboards and alerts. By tuning monitoring thresholds, you will enable proactive detection of issues and facilitate rapid troubleshooting. Your role will also involve conducting performance analysis and tuning for compute, storage, and network resources, contributing to capacity forecasting and supporting right-sizing efforts to optimize resource utilization.
Security, Patch & Resilience Operations
You will apply robust platform hardening measures, enforce patching schedules, and implement backup and recovery procedures to meet stringent security requirements. Your responsibilities include supporting disaster recovery exercises, verifying backup integrity, and deploying resilient patterns such as multi-AZ configurations and failover mechanisms to ensure high availability and business continuity.
Incident Response, Troubleshooting & L2/L3 Support
As the escalation point for complex incidents, you will triage, diagnose root causes, implement fixes or workarounds, and ensure incidents are resolved with documented Root Cause Analysis (RCA). You will collaborate closely with application teams, integration owners, and third-party vendors to resolve cross-stack faults and implement measures to prevent recurrence, thereby enhancing overall system reliability.
Reporting & Knowledge Sharing
You will provide periodic operational updates to the Section Head, covering deployments, incidents, risks, and improvement opportunities. Additionally, you will maintain comprehensive runbooks, configuration baselines, deployment playbooks, and change records while contributing to the team’s knowledge base and onboarding materials to foster continuous learning and operational excellence.
Policies, Processes and Procedures
You will conduct day-to-day activities in strict compliance with organizational policies and procedures while identifying opportunities for continuous improvement. Your role will involve evaluating systems and processes to incorporate leading industry practices, adapt to evolving business environments, reduce costs, and enhance productivity.
Qualifications & Requirements
- Experience: Minimum of 3 years in a related field or equivalent experience is required.
- Education: A bachelor’s degree in Engineering, Computer Science, Information Technology (IT), or an equivalent field is mandatory.
- Preferred Certifications: Infrastructure and architecture certifications such as Cisco CCNP, VMware VCP-DCV (for virtualization), ITIL, and CompTIA Network+, Security+, or A+. Fortinet NSE certifications are also advantageous.
يمكن أن يرتكب الذكاء الاصطناعي أخطاءً.
ℹ️ إخلاء مسؤولية توظيف:
موقع وظائف السعودية (ksajobshub.com) هو محرك بحث ومجمع لإعلانات الوظائف من المصادر والشركات الرسمية في المملكة العربية السعودية. نحن لا نتقاضى أي مبالغ مالية أو رسوم من الباحثين عن عمل، وتتم عمليات التقديم مباشرة عبر الانتقال للرابط الأصلي للجهة المعلنة.