Data & ML Ops

Salla
المدينة المنورة, المدينة المنورة دوام كامل
نشر: 1448/1/21 | 2026/07/06 ينتهي: 1448/2/22 | 2026/08/05 ✨ وصف بالذكاء الاصطناعي
تقدم للوظيفة الآن

الوصف الوظيفي

Senior Site Reliability Engineer (SRE)

We are seeking a Senior Site Reliability Engineer (SRE) to drive the design, scaling, and security of our rapidly expanding platform infrastructure. In this role, you will ensure the availability, performance, and cost-efficiency of our critical systems, from customer-facing applications and APIs to internal platforms and data services.

  • Work across diverse systems, ensuring reliability and efficiency at scale.
  • Leverage Kubernetes, observability, GitOps, automation, and cloud infrastructure to deliver a highly reliable and self-healing environment.
  • Collaborate with application, platform, and data teams to balance speed, stability, and cost-efficiency in production.

Key responsibilities include:

  • Designing, deploying, monitoring, and maintaining production workloads across Kubernetes clusters (EKS/AKS/GKE).
  • Building self-healing, auto-scaling systems to minimize manual intervention and ensure uptime.
  • Designing and operating reliable database and storage platforms within Kubernetes environments.
  • Implementing backup, disaster recovery, replication, and failover strategies to meet RPO/RTO targets.
  • Troubleshooting and recovering Kubernetes Persistent Volumes.
  • Optimizing storage performance and cost through multi-tier strategies, hot/cold data separation, and S3/offloading lifecycle policies.
  • Securing and scaling object storage platforms for high-throughput data pipelines.
  • Managing block storage and shared file systems for resilience and cost balance.
  • Collaborating with teams to optimize networking, ingress/egress traffic, and service mesh for secure communication.

Requirements:

  • Bachelor's degree in Computer Science, Engineering, or a related field, or equivalent work experience.
  • Proven experience in designing, deploying, and maintaining production workloads in Kubernetes environments.
  • Strong background in infrastructure reliability, automation, and delivery.
  • Expertise in observability, incident response, security, and performance optimization.

يمكن أن يرتكب الذكاء الاصطناعي أخطاءً.

المصدر: لينكد إن ↗ • 3 مشاهدة

وظائف مشابهة

تقدم للوظيفة الآن