Senior Site Reliability Engineer

Jobgether
السعودية, السعودية تدريب
نشر: 1448/4/10 | 2026/09/21 ينتهي: 1448/5/10 | 2026/10/21 ✨ وصف بالذكاء الاصطناعي
تقدم للوظيفة الآن مشاركة عبر واتساب

الوصف الوظيفي

Senior Site Reliability Engineer – Driving Operational Excellence in a Globally Distributed Environment

Join our high-performing infrastructure team as a Senior Site Reliability Engineer to ensure the seamless, high-performance operation of mission-critical production systems. In this pivotal role, you will take full ownership of system reliability, observability, and operational resilience while collaborating with cross-functional teams to enhance performance, security, and scalability. This is an opportunity to shape the future of our cloud-based infrastructure, reduce operational toil, and mentor engineers in best practices that drive efficiency and innovation.

As a key member of a globally distributed team, you will play a strategic role in maintaining the availability, latency, and correctness of our high-throughput platform—where every second of downtime or performance degradation directly impacts our customers. Your expertise will be instrumental in defining and enforcing reliability targets, optimizing monitoring and alerting systems, and proactively mitigating risks before they escalate into critical incidents.

Key Responsibilities

Your work will span a broad spectrum of technical and operational disciplines, including:

  • End-to-End Reliability Leadership: Own the reliability of critical production systems, encompassing instrumentation, performance under real-world traffic, and operational resilience. Define and maintain Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets to align engineering priorities with measurable reliability goals.
  • Observability and Alerting Optimization: Enhance alert quality by refining actionable signals, reducing noise, and addressing gaps in early detection of customer-impacting issues. Lead the improvement of monitoring and alerting frameworks to ensure proactive issue resolution.
  • Incident Response and Postmortem Excellence: Take a leading role in high-severity incident response, conducting thorough root-cause analyses, restoring service quickly, and driving actionable postmortems that prevent recurrence. Participate in on-call rotations while continuously refining runbooks and escalation pathways to minimize pager fatigue.
  • Resilient Infrastructure Design: Architect and operate secure, cost-efficient, and failure-resistant infrastructure with a focus on advanced failure modes—such as timeouts, retries, backpressure, load shedding, and graceful degradation. Implement blast-radius reduction strategies to minimize the impact of failures.
  • Performance and Capacity Planning: Conduct in-depth capacity and performance analysis through load testing, profiling, and saturation analysis. Proactively plan headroom to ensure scalability and prevent performance bottlenecks.
  • Automated Deployment and Change Safety: Improve change safety through progressive delivery, automated rollback mechanisms, and rigorous pre-production validation. Champion dependable deployment practices to minimize operational risk.
  • Infrastructure as Code (IaC) and Tooling: Manage infrastructure through code, primarily using Terraform, while ensuring consistency with broader service architecture. Develop production-grade software and developer-facing tooling that automates repetitive tasks and simplifies service ownership.
  • Chaos Engineering and Failure Testing: Design and execute controlled failure testing, game days, and chaos exercises to identify system vulnerabilities and establish safe operating limits.
  • Collaborative Product Engineering: Partner with product teams to ensure new and high-risk services are production-ready, covering capacity planning, failure modes, rollback strategies, and on-call handoffs.
  • Security and Mentorship: Apply a security-first mindset to all engineering work, identifying vulnerabilities in implementations and peer reviews. Mentor engineers through code reviews, pairing sessions, and design feedback to elevate team capabilities.

This role is ideal for an experienced engineer who thrives in a fast-paced, globally distributed environment and is passionate about operational excellence, automation, and continuous improvement. If you are driven by the challenge of maintaining large-scale, high-availability systems and enjoy mentoring peers while pushing the boundaries of infrastructure reliability, we invite you to apply.

يمكن أن يرتكب الذكاء الاصطناعي أخطاءً.

المصدر: لينكد إن ↗ • 4 مشاهدة

ℹ️ إخلاء مسؤولية توظيف:

موقع وظائف السعودية (ksajobshub.com) هو محرك بحث ومجمع لإعلانات الوظائف من المصادر والشركات الرسمية في المملكة العربية السعودية. نحن لا نتقاضى أي مبالغ مالية أو رسوم من الباحثين عن عمل، وتتم عمليات التقديم مباشرة عبر الانتقال للرابط الأصلي للجهة المعلنة.

وظائف مشابهة

تقدم للوظيفة الآن