Expert Site Reliability Engineer
الوصف الوظيفي
Expert Site Reliability Engineer – Elevate System Resilience and Operational Excellence
Join our high-performing team as an Expert Site Reliability Engineer, where you will play a pivotal role in ensuring the seamless, scalable, and resilient operation of our mission-critical technology infrastructure. In this strategic position, you will leverage cutting-edge reliability engineering principles, automation, and observability to drive operational excellence, minimize downtime, and enhance system performance. This is an opportunity to shape the future of our technology stack while fostering a culture of innovation, collaboration, and continuous improvement.
As an Expert Site Reliability Engineer, you will be responsible for defining and implementing best-in-class reliability engineering practices that underpin the availability, scalability, and fault tolerance of our core services. Your expertise will be instrumental in designing and optimizing systems that can withstand high demand, recover swiftly from failures, and adapt to evolving business needs. Whether it’s refining incident response workflows, automating repetitive tasks, or enhancing monitoring capabilities, your contributions will directly impact the reliability and efficiency of our technology ecosystem.
Key Responsibilities
Your role will encompass a broad spectrum of technical and leadership responsibilities, including:
- Define and Implement Advanced Reliability Practices: Lead the adoption of industry-leading Site Reliability Engineering (SRE) methodologies to ensure our systems meet the highest standards of availability, scalability, and resilience.
- Establish and Optimize Reliability Metrics: Define, monitor, and refine Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets to ensure alignment with business goals and operational excellence.
- Automate for Efficiency and Resilience: Design and implement automation solutions to reduce manual operational toil, streamline workflows, and enhance system robustness. Leverage Infrastructure as Code (IaC) and scripting to build scalable, repeatable, and maintainable environments.
- Enhance Observability and Alerting: Develop and refine monitoring, logging, and alerting systems to provide real-time visibility into system health, performance, and potential issues. Ensure alerts are actionable, reducing noise while improving incident detection.
- Lead Incident Response and Root Cause Analysis: Take ownership of complex production incidents, conduct thorough root cause analyses, and drive permanent corrective actions to prevent recurrence. Foster a proactive approach to incident management and continuous improvement.
- Drive Architectural and Engineering Improvements: Identify reliability risks and propose architectural enhancements that improve system availability, scalability, and disaster recovery capabilities. Collaborate with engineering teams to implement high-impact solutions.
- Advance Performance Engineering: Conduct capacity planning and performance tuning to ensure systems can handle peak loads and scale efficiently. Optimize resource utilization to balance cost and performance.
- Mentor and Guide Technical Teams: Provide expert guidance on SRE practices, automation, and reliability engineering to junior engineers and cross-functional teams. Foster a culture of shared knowledge and continuous learning.
- Reduce Operational Toil: Champion automation-driven approaches to minimize repetitive tasks, freeing up teams to focus on high-value initiatives that drive innovation and business growth.
This role is ideal for a seasoned professional who thrives in a fast-paced, tech-driven environment and is passionate about building reliable, scalable systems. Your expertise in cloud platforms, Kubernetes, and production-grade infrastructure will be critical in ensuring our technology services remain robust, secure, and aligned with business objectives.
Why This Role Matters
At its core, this position is about more than just fixing problems—it’s about preventing them. By embedding reliability engineering principles into our DNA, you will help us build systems that are not only resilient but also adaptable, efficient, and future-ready. Your work will directly contribute to the stability of our operations, the satisfaction of our users, and the success of our business. If you are a detail-oriented, analytical thinker with a passion for solving complex challenges, we invite you to be part of our mission to redefine operational excellence.
يمكن أن يرتكب الذكاء الاصطناعي أخطاءً.
ℹ️ إخلاء مسؤولية توظيف:
موقع وظائف السعودية (ksajobshub.com) هو محرك بحث ومجمع لإعلانات الوظائف من المصادر والشركات الرسمية في المملكة العربية السعودية. نحن لا نتقاضى أي مبالغ مالية أو رسوم من الباحثين عن عمل، وتتم عمليات التقديم مباشرة عبر الانتقال للرابط الأصلي للجهة المعلنة.