Site Reliability Engineer - Saudi Only
الوصف الوظيفي
Site Reliability Engineer – Ensuring Uninterrupted Excellence in a High-Growth AI Platform
At Lucidya, we are redefining customer experience (CX) intelligence with an AI-native platform that autonomously manages entire customer lifecycles—from initial engagement to retention and growth. Unlike traditional solutions that merely provide insights, our proprietary Natural Language Understanding (NLU) technology, built in-house and trained on millions of multilingual conversations, enables marketing, support, CX, and research teams to deliver hyper-personalized experiences. As we scale globally, the reliability, performance, and resilience of our infrastructure are not just operational priorities—they are the foundation of our mission to drive measurable improvements in customer satisfaction, retention, and lifetime value.
As a Site Reliability Engineer at Lucidya, you will play a pivotal role in ensuring our platform remains robust, scalable, and resilient under the most demanding conditions. This is not merely a technical role—it’s an opportunity to shape the reliability of a system that powers critical decision-making for our customers. If you thrive on solving complex infrastructure challenges, eliminating inefficiencies, and building systems that operate seamlessly—this is where your expertise will make a transformative impact.
Your Impact: Building a Reliable Future
In this role, you will transition from reactive problem-solving to proactive system design, ensuring that reliability is not an afterthought but a core principle of our infrastructure. Your responsibilities will span the entire spectrum of site reliability, from infrastructure design to incident response, automation, and continuous improvement. Here’s what success looks like:
- Design and Own Highly Available Systems: Architect and maintain cloud infrastructure that is fault-tolerant, scalable, and resilient, ensuring minimal downtime and optimal performance even as demand scales.
- Eliminate Single Points of Failure: Proactively identify and mitigate vulnerabilities before they escalate into incidents, safeguarding our platform’s stability.
- Optimize Cloud Environments: Manage and refine workloads across leading cloud providers (AWS, GCP, or Azure), leveraging Infrastructure as Code (Terraform) to standardize and scale infrastructure efficiently.
- Master Kubernetes in Production: Operate and scale Kubernetes clusters (EKS, GKE, etc.) with confidence, ensuring containerized workloads perform reliably at scale.
- Build Robust Observability: Implement and refine monitoring systems using tools like Prometheus, Grafana, Datadog, or ELK, while defining actionable alerting strategies that reduce noise and enhance responsiveness.
- Automate for Efficiency: Develop scripts and tooling to automate repetitive tasks, reducing manual intervention and accelerating deployment reliability.
- Drive Continuous Improvement: Collaborate with DevOps and engineering teams to address performance bottlenecks, refine CI/CD pipelines, and establish reliability best practices across the organization.
- Lead Incident Response: Respond to incidents with urgency, conduct thorough root cause analyses, and ensure lessons learned are institutionalized to prevent recurrence.
Your Journey: From Day One to Day 90
Your first 90 days will be a period of rapid integration and impact. By Day 30, you will:
- Develop a deep understanding of our infrastructure, systems, and operational workflows.
- Contribute actively to day-to-day operations, supported by your team.
- Begin identifying opportunities for automation and reliability enhancements.
By Day 90, you will:
- Independently manage infrastructure tasks and troubleshoot issues with confidence.
- Take ownership of specific infrastructure components, driving improvements in scalability and reliability.
- Actively shape the reliability culture within the organization, ensuring a proactive and resilient approach to system operations.
Who You Are: The Ideal Site Reliability Engineer
This role is designed for engineers who:
- Possess 3+ years of experience in Site Reliability Engineering (SRE), DevOps, or infrastructure engineering, with a proven track record of designing and maintaining scalable, high-availability systems.
- Are passionate about automation and efficiency, leveraging tools and practices to eliminate manual processes and reduce operational overhead.
- Have a deep understanding of cloud-native technologies, including Kubernetes, Terraform, and major cloud platforms (AWS, GCP, or Azure).
- Excel in observability and monitoring, with hands-on experience in tools like Prometheus, Grafana, or ELK, and a focus on meaningful alerting strategies.
- Are collaborative problem-solvers, thriving in cross-functional environments and contributing to both technical and organizational growth.
- Share our commitment to excellence and innovation, driving reliability not just as a technical requirement, but as a competitive advantage.
Join Lucidya and be part of a team that is redefining the future of customer experience intelligence. Your expertise in site reliability will ensure our platform remains the backbone of our customers’ success—today and as we scale tomorrow.
يمكن أن يرتكب الذكاء الاصطناعي أخطاءً.
ℹ️ إخلاء مسؤولية توظيف:
موقع وظائف السعودية (ksajobshub.com) هو محرك بحث ومجمع لإعلانات الوظائف من المصادر والشركات الرسمية في المملكة العربية السعودية. نحن لا نتقاضى أي مبالغ مالية أو رسوم من الباحثين عن عمل، وتتم عمليات التقديم مباشرة عبر الانتقال للرابط الأصلي للجهة المعلنة.