Site Reliability Engineer

TestCrew | Quality Engineering & Software Testing
الأحساء, الأحساء دوام كامل
نشر: 1448/3/27 | 2026/09/09 ينتهي: 1448/4/28 | 2026/10/09 ✨ وصف بالذكاء الاصطناعي
تقدم للوظيفة الآن

الوصف الوظيفي

Site Reliability Engineer – Enterprise Engagement in Al Ahsa

Join TestCrew as a **Site Reliability Engineer (SRE)** to drive operational excellence and reliability for mission-critical IT systems in a dynamic enterprise environment. This is an opportunity to make a tangible impact by ensuring seamless uptime, optimizing performance, and fostering collaboration across development, operations, and security teams. As an on-site SRE in Al Ahsa, you will play a pivotal role in maintaining high availability, automating workflows, and implementing best-in-class observability and incident response strategies.

In this hands-on role, you will bridge the gap between engineering and operations to deliver scalable, resilient, and secure solutions. Your expertise in monitoring, automation, and troubleshooting will be instrumental in reducing downtime, improving system efficiency, and supporting continuous service innovation. This is a chance to work in a fast-paced, collaborative environment where your technical leadership directly enhances business continuity and operational resilience.

Key Responsibilities

As the Site Reliability Engineer, your responsibilities will include:

  • Ensuring High Availability and Performance: Monitor and manage enterprise-grade infrastructure, applications, and services to guarantee optimal uptime, stability, and performance while adhering to strict SLAs.
  • Proactive Incident Management: Design and implement robust monitoring and observability platforms (e.g., Grafana, Prometheus, Datadog) to detect anomalies, analyze trends, and respond swiftly to incidents with disciplined root cause analysis (RCA).
  • Automation and Efficiency: Drive operational excellence by automating repetitive tasks, deployments, and maintenance activities using scripting (Python, Bash, PowerShell) and Infrastructure as Code (IaC) tools like Terraform and Ansible.
  • CI/CD and Release Management: Support and optimize CI/CD pipelines (Jenkins, GitLab CI, Azure DevOps) to enable seamless, reliable software deployments while maintaining version control and release governance.
  • Security and Compliance: Implement security best practices, including system hardening, patch management, and secure operational procedures to mitigate risks and ensure compliance with enterprise standards.
  • Collaborative Problem-Solving: Partner closely with development, infrastructure, security, and IT operations teams to resolve complex technical challenges and drive cross-functional improvements.
  • Documentation and Knowledge Sharing: Maintain comprehensive operational documentation, runbooks, and incident reports to foster transparency and support continuous learning within the team.
  • Incident Response and Improvement: Participate in major incident management, post-mortem reviews, and iterative process enhancements to strengthen operational resilience.
  • On-Call Support: Provide round-the-clock on-call coverage and participate in shift rotations to ensure 24/7 system availability and rapid incident resolution.

Why This Role Matters

This is more than a technical role—it’s an opportunity to shape the reliability and scalability of critical systems for an enterprise client. Your work will directly contribute to minimizing downtime, enhancing user experience, and supporting business continuity. By leveraging your expertise in SRE principles, automation, and observability, you will play a key role in transforming operational workflows and setting new standards for efficiency and security.

Qualifications and Skills

To excel in this role, you will bring:

  • Education: A Bachelor’s degree in Computer Science, Information Technology, Engineering, or a related field.
  • Experience: 2–5 years of hands-on experience in Site Reliability Engineering (SRE), DevOps, Infrastructure Operations, or Production Support, with a focus on enterprise environments.
  • Monitoring and Observability: Proven experience managing platforms such as Grafana, Prometheus, Datadog, Instana, or Zabbix, with a strong ability to design dashboards, alerts, and performance metrics.
  • Troubleshooting and RCA: Exceptional skills in diagnosing and resolving complex issues across infrastructure, applications, and networking layers.
  • Automation and Scripting: Proficiency in scripting languages (Python, Bash, PowerShell) and automation tools (Ansible, Terraform) to streamline operational tasks.
  • CI/CD Expertise: Practical knowledge of CI/CD pipelines and tools (Jenkins, GitLab CI, Azure DevOps) to support efficient and reliable deployments.
  • Security and Compliance: A solid understanding of cybersecurity best practices, including system hardening, access control, vulnerability management, and patch management.
  • Language Proficiency: Fluency in both Arabic and English, with excellent written and verbal communication skills to facilitate collaboration across diverse teams.
  • Preferred Qualifications:
    - ITIL Foundation Certification
    - Experience with cloud platforms (AWS, Azure, GCP)
    - Knowledge of containerization and orchestration (Docker, Kubernetes)
    - Background in supporting large-scale enterprise or government environments

If you are passionate about reliability engineering, thrive in collaborative environments, and are eager to make a meaningful impact in a high-stakes enterprise setting, we invite you to apply. Join TestCrew and help redefine operational excellence in Al Ahsa.

يمكن أن يرتكب الذكاء الاصطناعي أخطاءً.

المصدر: لينكد إن ↗ • 4 مشاهدة

ℹ️ إخلاء مسؤولية توظيف:

موقع وظائف السعودية (ksajobshub.com) هو محرك بحث ومجمع لإعلانات الوظائف من المصادر والشركات الرسمية في المملكة العربية السعودية. نحن لا نتقاضى أي مبالغ مالية أو رسوم من الباحثين عن عمل، وتتم عمليات التقديم مباشرة عبر الانتقال للرابط الأصلي للجهة المعلنة.

وظائف مشابهة

تقدم للوظيفة الآن