Lead Site Reliability Engineer - Imunify Reliability Platform
الوصف الوظيفي
Lead Site Reliability Engineer – Imunify Reliability Platform
Join a cutting-edge security-focused organization as the **Lead Site Reliability Engineer** for the Imunify Reliability Platform—a groundbreaking opportunity to architect and refine the reliability framework for a large-scale, mission-critical product. In this **greenfield leadership role**, you will define operational excellence from the ground up, ensuring seamless performance, security, and resilience across approximately **70 critical components**, spanning cloud services and customer-hosted agents. This is your chance to shape the future of reliability engineering in a high-impact, security-centric environment while driving measurable outcomes that directly enhance customer trust and system integrity.
As the **first-of-its-kind SRE leader**, you will collaborate closely with engineering leads and senior stakeholders to establish industry-leading **Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets**, ensuring proactive detection of silent security-control degradation before it escalates into widespread disruptions. Your work will define the **telemetry, monitoring, and alerting infrastructure**, fostering a culture of reliability and operational accountability. This is not just about maintaining systems—it’s about **building a self-healing, observable, and resilient platform** that adapts to evolving threats and scalability demands.
Key Responsibilities
Your role will encompass a broad spectrum of strategic and technical responsibilities, including but not limited to:
- Define and Implement a Robust Reliability Framework: Establish **meaningful SLIs and SLOs** for all 70+ product components, working collaboratively with engineering squads to define ownership, measurement tiers, and error budgets that align with business and security objectives.
- Develop a Comprehensive Reliability Taxonomy: Create a structured taxonomy covering **service availability, latency, fleet reachability, security-control efficacy, artifact delivery, and telemetry pipeline health**—ensuring every metric is independently verifiable and resistant to manipulation by the failures it monitors.
- Architect a Scalable Telemetry Pipeline: Design and build a **high-performance telemetry collection system** for both cloud services and customer-hosted agents, balancing **push-based sampling, privacy constraints, data quality, and cardinality** to support real-time observability without overhead.
- Enhance Instrumentation Across Critical Languages: Collaborate with product engineering teams to extend **Prometheus/OpenMetrics-based instrumentation** in **Python, Go, and Rust**, ensuring comprehensive visibility into system behavior and performance.
- Optimize Observability Platforms: Consolidate existing dashboards, queries, and reporting mechanisms into a **leaner, more reliable observability stack**, retiring redundant tooling to eliminate operational noise while maximizing actionable insights.
- Implement SLO-Driven Alerting: Develop **symptom-based, multi-window alerting** with **burn-rate principles**, ensuring alerts are **clear, actionable, and classified** by severity—with defined owners, documented failure modes, and runbooks for every production incident.
- Foster Alert Quality and Ownership: Establish **ongoing alert-quality practices**, including periodic reviews, measurable alert actionability, and deliberate removal of unnecessary alerts to reduce alert fatigue and improve incident response.
- Build a Machine-Readable Escalation Model: Create a **scalable ownership and escalation framework** that routes incidents efficiently across **global time zones**, ensuring follow-the-sun support, clear severity definitions, and seamless handoffs.
- Strengthen Incident Response and Postmortems: Implement **blameless postmortem practices** with reliable timelines, ownership tracking, and actionable follow-through to prevent recurrence, while coaching engineering squads to **embrace operational ownership** rather than relying on centralized support.
- Drive Measurable Reliability Outcomes: Deliver **tangible results within the first year**, including full SLI ownership, comprehensive production telemetry, tiered alerting adoption, and a **significant reduction in silent security-control degradation detection time**—directly improving customer trust and system resilience.
Why This Opportunity Stands Out
This is more than a job—it’s a **chance to pioneer reliability engineering in a security-first environment**. You will:
- Work in a **remote-first, async-driven culture** where **technical excellence and measurable impact** take precedence over process overhead.
- Collaborate with **top-tier engineering leads** to shape the **reliability culture** of a major security product.
- Leverage **cutting-edge technologies** like **Prometheus, Grafana, Alertmanager, ClickHouse, and distributed systems debugging** to solve complex, real-world challenges.
- Contribute to a **platform that proactively detects and mitigates security risks**, ensuring customers remain protected in an evolving threat landscape.
- Be part of a **high-performing team** that values **ownership, accountability, and continuous improvement**—where every decision drives tangible business value.
If you are a **visionary SRE with a passion for building resilient, observable systems**, this is your opportunity to **leave a lasting impact** in one of the most critical areas of modern cybersecurity. Join us in defining the future of reliability—where **technical depth meets strategic vision**.
يمكن أن يرتكب الذكاء الاصطناعي أخطاءً.
ℹ️ إخلاء مسؤولية توظيف:
موقع وظائف السعودية (ksajobshub.com) هو محرك بحث ومجمع لإعلانات الوظائف من المصادر والشركات الرسمية في المملكة العربية السعودية. نحن لا نتقاضى أي مبالغ مالية أو رسوم من الباحثين عن عمل، وتتم عمليات التقديم مباشرة عبر الانتقال للرابط الأصلي للجهة المعلنة.