Site Reliability Engineering Officer

Takamol Holding
الرياض, الرياض دوام كامل
نشر: 1448/1/20 | 2026/07/05 ينتهي: 1448/2/21 | 2026/08/04 ✨ وصف بالذكاء الاصطناعي
تقدم للوظيفة الآن

الوصف الوظيفي

Site Reliability Engineering Officer

We are seeking a dedicated Site Reliability Engineering Officer to join our dynamic team. In this role, you will provide critical support for application incidents across our digital platforms, working collaboratively with Platform Engineering, Application Development, and customer support teams to ensure timely resolution according to established Service Level Agreements (SLAs) and escalation procedures.

Your key responsibilities will include:

  • Operating and monitoring the Elastic Observability stack, including Elasticsearch cluster health, Kibana, Fleet Server, APM Server, and Elastic Agent, deployed and managed via ECK on OKE.
  • Assisting with day-to-day Elasticsearch operations such as index lifecycle management (ILM), snapshot lifecycle management (SLM), data tier housekeeping, and capacity monitoring.
  • Troubleshooting telemetry ingestion issues across logs, metrics, traces, and synthetic monitors to ensure consistent data collection from all platforms.
  • Maintaining and updating Kibana dashboards, alerting rules, and saved objects under the guidance of the SRE Manager.
  • Performing root cause analysis and participating in blameless post-incident reviews to improve system reliability and reduce recurrence.
  • Collaborating with Platform Engineering to automate repetitive tasks, improve deployment pipelines, and enhance observability coverage using Terraform, Helm charts, and scripting.
  • Developing and maintaining support documentation, runbooks, and knowledge base articles aligned to standardized incident response procedures.
  • Managing and prioritizing incidents and requests via the ticketing system (Jira/ServiceNow), ensuring all incidents, requests, and resolutions are documented in the service management system.
  • Participating in an on-call rotation and helping reduce operational toil through automation and tooling.
  • Monitoring and reporting on key performance metrics related to incident management, including mean time to detect (MTTD) and mean time to resolve (MTTR).
  • Collaborating with cross-functional teams and vendor partners to improve overall system reliability, observability maturity, and security posture.

To be successful in this role, you should possess:

  • A Bachelor’s degree in Computer Science, IT, Engineering, or a related field (or equivalent experience).
  • 1–3 years of experience in IT operations, system administration, application support, DevOps, or SRE.
  • Familiarity with Observability tools such as Elastic Stack (Elasticsearch, Kibana, etc.), including basic querying and dashboard usage.
  • Knowledge of Linux systems and scripting (Bash, Python, or Go).
  • Understanding of monitoring, logging, and alerting concepts.
  • Experience with ITSM tools (ServiceNow, Jira, Zendesk) and ITIL practices.
  • A strong grasp of incident, problem, and change management.
  • Basic experience with cloud native environments and containers such as Docker and Kubernetes.
  • Strong critical thinking, troubleshooting, and communication skills.

يمكن أن يرتكب الذكاء الاصطناعي أخطاءً.

المصدر: لينكد إن ↗ • 4 مشاهدة

وظائف مشابهة

تقدم للوظيفة الآن