Senior DevOps Engineer

SmartChoice International GCC
الرياض, الرياض دوام كامل
نشر: 1448/3/27 | 2026/09/09 ينتهي: 1448/4/28 | 2026/10/09 ✨ وصف بالذكاء الاصطناعي
تقدم للوظيفة الآن مشاركة عبر واتساب

الوصف الوظيفي

Senior DevOps Engineer – Cloud Infrastructure & Platform Engineering

Join a forward-thinking technology organization in Riyadh as a **Senior DevOps Engineer**, where you will architect, deploy, and optimize the scalable, secure, and high-performance infrastructure powering a rapidly evolving platform. This hands-on role demands expertise in cloud-native technologies, automation, and operational excellence to drive efficiency, reliability, and cost optimization across production environments. You will collaborate closely with engineering teams to establish best-in-class operational standards, ensuring seamless deployment pipelines, resilient infrastructure, and robust observability—all while fostering a culture of continuous improvement.

As a key contributor to the platform engineering function, you will design and implement infrastructure-as-code solutions, automate CI/CD workflows, and manage containerized workloads with precision. Your leadership in incident response, security hardening, and performance tuning will elevate the organization’s operational maturity, enabling teams to deliver services with confidence and agility.

Key Responsibilities

  • Infrastructure Design & Automation: Architect and maintain scalable, reproducible cloud infrastructure using Terraform/OpenTofu, ensuring modularity, reusability, and state management best practices. Leverage infrastructure-as-code to streamline deployments and reduce operational overhead.
  • CI/CD & Deployment Excellence: Own end-to-end CI/CD pipelines, integrating build, test, security scanning, and production deployment stages. Implement robust rollout strategies, automated rollbacks, and disaster recovery mechanisms to minimize downtime and risk.
  • Container & Kubernetes Mastery: Manage containerized applications across Kubernetes (EKS) and hybrid cloud environments, optimizing resource allocation, scaling, and observability. Design Helm charts and GitOps workflows (e.g., ArgoCD/Flux) to ensure declarative, version-controlled deployments.
  • Security & Compliance: Enforce least-privilege access, secrets management (Vault/AWS Secrets Manager), and identity governance. Harden infrastructure against threats while ensuring compliance with organizational security policies.
  • Database & High Availability: Operate production-grade databases (PostgreSQL, RDS/Aurora, Redis), configuring high availability, replication, backups, and failover strategies to guarantee data integrity and minimal latency.
  • Observability & Reliability: Establish comprehensive monitoring, alerting, and metrics using Prometheus, Grafana, CloudWatch, and OpenTelemetry. Proactively identify bottlenecks, optimize performance, and reduce mean time to resolution (MTTR) through data-driven insights.
  • Incident Response & Resilience: Lead post-mortems, analyze root causes, and implement corrective actions to prevent recurrence. Champion a culture of reliability engineering, ensuring systems recover gracefully from failures.
  • Cost & Performance Optimization: Analyze infrastructure utilization, right-size resources, and implement autoscaling strategies to balance performance with cost efficiency—particularly for high-compute workloads like AI/ML.
  • Platform Enablement: Develop reusable platform capabilities (e.g., API gateways, service meshes) to empower engineering teams with self-service deployment tools while maintaining governance and security.

Technical Environment & Expertise

This role demands deep expertise in a cloud-first architecture, with a strong preference for AWS across compute, networking, IAM, and storage. Proficiency in the following technologies is essential:

  • Cloud & Infrastructure: AWS (EC2, VPC, IAM, S3, EBS), Terraform/OpenTofu, Ansible, and Bash/Python scripting for automation.
  • Containers & Orchestration: Docker, Kubernetes (EKS), Helm, and GitOps tools like ArgoCD/Flux.
  • CI/CD & DevOps: GitHub Actions (or equivalent), pipeline orchestration, and immutable infrastructure practices.
  • Security & Secrets Management: Vault, AWS Secrets Manager, and IAM policies for least-privilege access.
  • Databases & Caching: PostgreSQL, RDS/Aurora, and Redis for high-performance, scalable data layers.
  • Observability: Prometheus, Grafana, CloudWatch, and OpenTelemetry for metrics, logs, and tracing.
  • High-Compute Workloads (Desirable): Experience with AI/ML infrastructure, including GPU scheduling, model serving (vLLM, Triton, KServe, Ray Serve), and cost-optimized autoscaling for SageMaker or Bedrock deployments.

Why This Role Matters

You are not just an engineer—you are a strategic partner in building a platform that scales with the business. Your work will directly impact the reliability, security, and cost efficiency of mission-critical systems, enabling the organization to innovate faster while maintaining enterprise-grade resilience. If you thrive in a fast-paced, collaborative environment and are passionate about turning complexity into simplicity, this is your opportunity to make a lasting technical impact.

Join us to shape the future of infrastructure engineering in a dynamic, high-growth setting where your expertise will drive both technical and business outcomes.

يمكن أن يرتكب الذكاء الاصطناعي أخطاءً.

المصدر: لينكد إن ↗ • 23 مشاهدة

ℹ️ إخلاء مسؤولية توظيف:

موقع وظائف السعودية (ksajobshub.com) هو محرك بحث ومجمع لإعلانات الوظائف من المصادر والشركات الرسمية في المملكة العربية السعودية. نحن لا نتقاضى أي مبالغ مالية أو رسوم من الباحثين عن عمل، وتتم عمليات التقديم مباشرة عبر الانتقال للرابط الأصلي للجهة المعلنة.

وظائف مشابهة

تقدم للوظيفة الآن