Staff Site Reliability Engineer

Jobgether
السعودية, السعودية تدريب تدريب / بدون خبرة
نشر: 1448/4/5 | 2026/09/16 ينتهي: 1448/5/5 | 2026/10/16 ✨ وصف بالذكاء الاصطناعي
تقدم للوظيفة الآن مشاركة عبر واتساب

الوصف الوظيفي

Staff Site Reliability Engineer – Leadership Role in Global Reliability Engineering

We are seeking a **Staff Site Reliability Engineer** to join our fully remote, globally distributed engineering organization as the inaugural SRE leader. This high-impact role demands a strategic mindset, hands-on technical expertise, and a passion for fostering a culture of reliability across critical production systems. As the first dedicated SRE, you will establish foundational reliability practices, drive measurable improvements in system resilience, and empower engineering teams to own operational excellence.

In this influential position, you will bridge the gap between engineering, infrastructure, and product teams to embed **Site Reliability Engineering (SRE) principles** into the organization’s DNA. Your work will shape how reliability is defined, measured, and prioritized—ensuring that systems are not only robust but also aligned with business objectives. With substantial autonomy, you will define industry-leading standards, mentor engineers, and build scalable practices that evolve alongside the company’s growth.

Key Responsibilities

Your role will encompass a broad spectrum of responsibilities, ensuring reliability is a core pillar of our engineering culture:

  • Define and Enforce Reliability Metrics: Establish **Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets** to quantify reliability and guide engineering decisions. Ensure these metrics are visible, actionable, and integrated into performance reviews.
  • Lead Incident Management: Strengthen the entire incident lifecycle—from detection and response to communication, postmortems, and follow-up actions. Optimize alerting, escalation pathways, and runbooks to minimize downtime and improve operational readiness.
  • Advocate for Proactive Reliability: Introduce **deliberate failure testing, game days, and chaos engineering** to identify system vulnerabilities before they impact production. Collaborate with teams to validate operational limits and rollback strategies.
  • Drive Cultural Shift: Coach **Staff and Lead Engineers** to adopt SRE principles, fostering a distributed ownership model where reliability is a shared responsibility. Develop lightweight operational standards for production readiness, on-call practices, and change safety.
  • Leverage AI for Operational Excellence: Explore and implement AI-driven solutions for **incident investigation, observability, and runbook automation**. Structure operational data to ensure both human engineers and AI systems can safely interpret and act on production signals.
  • Hands-On Problem Solving: Engage directly in production incidents, developing tooling, dashboards, and automation to prevent recurrence. Contribute directly to code and infrastructure improvements, ensuring reliability enhancements are actionable and scalable.
  • Shape System Design: Partner with architects and technical leads to embed **resilience and failure tolerance** into system design, ensuring reliability is considered from the ground up rather than retroactively applied.

Why This Role Matters

This is more than an engineering position—it’s an opportunity to **define how reliability is perceived and practiced** within a rapidly growing organization. Your influence will extend beyond technical implementations, shaping the mindset of engineers and leaders alike. If you thrive in a high-impact, collaborative environment where your expertise directly impacts global systems, we invite you to take the next step in your career.

Requirements

To excel in this role, you will bring:

  • 10+ years of engineering experience**, with at least **3 years in SRE, production engineering, or a Staff Engineer role** focused on reliability at a platform or organizational level.
  • Proven expertise in designing and implementing SLIs, SLOs, and error budgets**, with a track record of driving adoption across engineering and product teams.
  • Deep incident leadership experience**, including managing high-severity incidents with clarity, accountability, and measurable improvements.
  • A hands-on approach to reliability**, combining technical depth with the ability to mentor and influence cross-functional teams.
  • Experience leveraging AI and automation** to enhance observability, incident response, and operational tooling.
  • Strong communication skills** to translate technical reliability concepts into actionable strategies for non-technical stakeholders.

This role is ideal for engineers who are not only technically proficient but also passionate about **building scalable, resilient systems** and fostering a culture where reliability is everyone’s responsibility.

يمكن أن يرتكب الذكاء الاصطناعي أخطاءً.

المصدر: لينكد إن ↗ • 6 مشاهدة

ℹ️ إخلاء مسؤولية توظيف:

موقع وظائف السعودية (ksajobshub.com) هو محرك بحث ومجمع لإعلانات الوظائف من المصادر والشركات الرسمية في المملكة العربية السعودية. نحن لا نتقاضى أي مبالغ مالية أو رسوم من الباحثين عن عمل، وتتم عمليات التقديم مباشرة عبر الانتقال للرابط الأصلي للجهة المعلنة.

وظائف مشابهة