Site Reliability Engineer
الوصف الوظيفي
About Lucidya
Lucidya is a cutting-edge AI-native platform designed for customer experience (CX) intelligence, empowering organizations to autonomously manage entire customer lifecycles—from initial engagement to retention and growth. Unlike conventional platforms that merely provide insights without action, Lucidya bridges the gap with proprietary Natural Language Understanding (NLU) technology developed in-house and trained on millions of multilingual conversations. This enables marketing, support, CX, and research teams to deliver hyper-personalized experiences that drive measurable improvements in customer satisfaction, retention, and lifetime value.
As Lucidya scales its global presence, the reliability, performance, and resilience of our infrastructure have become mission-critical to our operations and customer success. This is where you come in. As a Site Reliability Engineer, you will play a pivotal role in ensuring our platform remains robust, scalable, and resilient, enabling our customers to make data-driven decisions without interruption.
Why This Role Matters
At Lucidya, our platform processes vast volumes of real-time customer data. Any downtime, latency, or instability directly impacts our customers' ability to serve their users effectively. This role is designed to prevent such scenarios by ensuring our infrastructure operates flawlessly, scales seamlessly, and adapts to evolving demands. You will not only react to issues but anticipate them, design preventive systems, and automate solutions to eliminate inefficiencies entirely.
If you are passionate about solving complex infrastructure challenges, optimizing performance, and building systems that operate with minimal manual intervention, this is the role for you. You will be at the heart of our platform’s stability, driving reliability as a core competency and enabling our teams to focus on innovation and growth.
What You’ll Do
In this role, you will be responsible for outcomes that directly impact our platform’s reliability and scalability. Your responsibilities will include:
- Ensuring Reliability as a Default: Design and maintain infrastructure that is highly available, fault-tolerant, and scalable to meet the demands of our growing customer base.
- Proactively Eliminating Risks: Identify and address single points of failure before they escalate into incidents, ensuring our systems remain resilient under pressure.
- Optimizing Cloud Environments: Manage and continuously improve workloads across major cloud platforms, including AWS, GCP, or Azure, to balance performance and cost efficiency.
- Infrastructure as Code (IaC): Use tools like Terraform to standardize and scale infrastructure deployments, ensuring consistency and reducing manual errors.
- Kubernetes Operations: Run and improve Kubernetes in production, operating and scaling clusters (such as EKS or GKE) with confidence while ensuring containerized workloads perform reliably at scale.
- Enhancing Observability: Implement and refine monitoring systems using tools like Prometheus, Grafana, Datadog, or ELK to gain deep insights into system performance and health.
- Meaningful Alerting: Define alerting mechanisms that are actionable and not overly noisy, ensuring teams respond to critical issues promptly and effectively.
- Incident Response and Root Cause Analysis: Lead incident response efforts, conduct thorough root cause analyses, and implement measures to prevent recurrence, fostering a culture of continuous improvement.
- Automating Operational Work: Write scripts and build tooling to eliminate repetitive tasks, promoting a culture where manual work is a temporary state rather than the norm.
- Collaborating Across Teams: Work closely with DevOps and engineering teams to resolve performance bottlenecks, contribute to CI/CD improvements, and help shape reliability best practices across the organization.
What Success Looks Like (First 90 Days)
First 30 Days:
- Develop a deep understanding of our infrastructure, systems, and operational workflows.
- Begin contributing to day-to-day operations with guidance from the team.
- Identify opportunities for automation and reliability improvements.
By 90 Days:
- Independently manage infrastructure tasks and troubleshoot issues with confidence.
- Actively contribute to reliability and scalability improvements across the platform.
- Take ownership of specific infrastructure components and drive enhancements.
Requirements
Who You Are: To excel in this role, you should have:
- At least three years of experience in Site Reliability Engineering (SRE), DevOps, or infrastructure engineering.
- A strong foundation in cloud platforms such as AWS, GCP, or Azure, with hands-on experience managing and optimizing cloud workloads.
- Proficiency in Infrastructure as Code (IaC) tools like Terraform, with a focus on standardization and scalability.
- Experience operating and scaling Kubernetes clusters in production environments (e.g., EKS, GKE).
- Expertise in monitoring and observability tools such as Prometheus, Grafana, Datadog, or ELK, with a keen eye for meaningful alerting.
- A proactive mindset with a passion for automating manual processes and eliminating inefficiencies.
- Strong problem-solving skills and the ability to troubleshoot complex infrastructure issues under pressure.
- Excellent collaboration skills, with the ability to work effectively with cross-functional teams to drive reliability improvements.
- A commitment to fostering a culture of reliability, where manual work is minimized, and systems are designed for resilience.
يمكن أن يرتكب الذكاء الاصطناعي أخطاءً.
ℹ️ إخلاء مسؤولية توظيف:
موقع وظائف السعودية (ksajobshub.com) هو محرك بحث ومجمع لإعلانات الوظائف من المصادر والشركات الرسمية في المملكة العربية السعودية. نحن لا نتقاضى أي مبالغ مالية أو رسوم من الباحثين عن عمل، وتتم عمليات التقديم مباشرة عبر الانتقال للرابط الأصلي للجهة المعلنة.