Automation & Observability Engineer
الوصف الوظيفي
Automation & Observability Engineer – Drive Efficiency and Reliability in Banking Infrastructure
Join our dynamic team as an **Automation & Observability Engineer**, where you’ll play a pivotal role in enhancing operational resilience, automating critical workflows, and ensuring seamless disaster recovery for mission-critical banking applications. In this strategic position, you’ll leverage cutting-edge tools like **Ansible Automation Platform, Python, and observability frameworks (ELK & Grafana)** to streamline infrastructure management, reduce manual intervention, and bolster system reliability. Your expertise will directly contribute to optimizing performance, minimizing downtime, and aligning technology with regulatory and security standards.
As an Automation & Observability Engineer, you’ll be at the forefront of innovation, transforming complex processes into scalable, automated solutions. Your work will bridge the gap between development, operations, and security, ensuring seamless integration across hybrid environments—spanning **Windows, AIX, and RHEL**—while maintaining compliance and performance excellence.
Key Responsibilities
Your role will encompass a broad spectrum of responsibilities, ensuring end-to-end automation and observability across our infrastructure:
- Automation Development & Maintenance: Design, develop, and maintain **Ansible playbooks, roles, and collections** to automate infrastructure provisioning, configuration management, and application deployments across **Windows Server, AIX, and RHEL**, including databases (MSSQL, Oracle, DB2, Teradata) and middleware (WAS, MQ). Utilize **Jinja2 templating** for dynamic configuration and **Ansible Vault** for secure credential management, integrating seamlessly with **Git, Bitbucket, and CI/CD pipelines**.
- Disaster Recovery & Orchestration: Develop and validate **disaster recovery job templates** and consistency checks to ensure business continuity. Troubleshoot and refine automation workflows to enhance reliability and reduce recovery time objectives (RTOs).
- Observability & Dashboarding: Design and implement **ELK (Elasticsearch/Logstash/Kibana) and Grafana dashboards** to monitor system health, performance metrics, and application logs. Leverage these insights to proactively identify and resolve issues, ensuring operational transparency and data-driven decision-making.
- Scripting & Tooling: Develop **Python CLI tools, API clients, and data parsers**, alongside **Bash/ksh and PowerShell scripts** for Windows automation. Implement robust **logging, error handling, and unit testing** frameworks to ensure script reliability. Package and reuse scripts using **virtual environments (venv), pip, and modular design principles** for maintainability.
- Integration & API Development: Build **REST API integrations and webhooks** to connect automation workflows with **ITSM platforms (e.g., ServiceNow)**, enabling seamless incident, change, and problem management. Collaborate with security teams to integrate **SIEM feeds, audit requirements, and regulatory retention policies** into observability pipelines.
- Cross-Functional Collaboration: Partner with security, development, and operations teams to align automation strategies with **compliance, scalability, and performance goals**. Contribute to documentation and best practices for **Git workflows, CI/CD pipelines, and change management** to ensure consistency and reproducibility.
Why This Role Matters
In today’s fast-paced financial landscape, automation and observability are non-negotiable pillars of operational success. Your work will directly impact:
- Operational Resilience: Reduce human error and accelerate recovery from failures, ensuring high availability for banking services.
- Efficiency Gains: Automate repetitive tasks, freeing up teams to focus on strategic initiatives and innovation.
- Regulatory Compliance: Align infrastructure and automation practices with industry standards, enhancing audit readiness and risk mitigation.
- Data-Driven Decision Making: Provide actionable insights through observability tools, enabling proactive issue resolution and performance optimization.
Qualifications & Technical Expertise
To thrive in this role, you’ll bring a blend of hands-on technical skills and strategic thinking:
- Experience: 4–8+ years in **Infrastructure, Platform Engineering, SRE, or DevOps**, with a proven track record of automating complex environments.
- Ansible Mastery: Deep expertise in **Ansible playbooks, roles, and collections**, including advanced use of **Jinja2 templating, Ansible Vault, and integration with version control systems**.
- Scripting Proficiency: Strong proficiency in **Python, Bash/ksh, and PowerShell**, with experience developing **production-grade scripts** for automation, monitoring, and data processing.
- Operating System Administration: Hands-on experience administering **Windows Server, AIX, and RHEL**, including **AD integration, LPAR management, systemd, SELinux, and kernel tuning**.
- Observability & Monitoring: Extensive experience with **ELK Stack (Elasticsearch/Logstash/Kibana) and Grafana**, including **log ingestion, parsing, and dashboard customization** for real-time monitoring.
- API & Integration Skills: Experience with **REST APIs, webhooks, and data integrations**, with fluency in **JSON/YAML** and integration frameworks like **ServiceNow, GTM, and Infoblox**.
- CI/CD & DevOps Practices: Familiarity with **Git workflows, CI/CD pipelines, and documentation best practices**, ensuring automated workflows are scalable and maintainable.
- Troubleshooting Acumen: Strong diagnostic skills across **networking, OS layers, and application logs**, with a focus on root-cause analysis and resolution.
Preferred Qualifications
While not mandatory, the following will further elevate your impact in this role:
- Certifications: **RHCSA/RHCE, Certified Ansible Engineer, Red Hat Ansible Automation, or Microsoft Windows Server/PowerShell certifications**.
- ITSM & Monitoring Tools: Experience with **ServiceNow (events/incidents/CMDB), Satellite, BigFix, Tanium, AppDynamics, Dynatrace, or Prometheus**.
- Cloud & Containers: Familiarity with **Docker, Kubernetes, and Oracle Cloud Infrastructure (OCI)** is a significant advantage.
- Advanced Automation Platforms: Knowledge of **vRO (vRealize Orchestrator), Airflow, Control-M, and vSphere** for orchestration and workflow automation.
- AIX & RHEL Deep Dive: Advanced skills in **AIX LPAR management, smitty, nim concepts, LVM, and ksh scripting**, as well as **RHEL systemd, SELinux, firewalld, and dnf/yum management**.
- Windows PowerShell DSC: Experience with **PowerShell Desired State Configuration (DSC)** for infrastructure-as-code in Windows environments.
If you’re passionate about **automation, observability, and driving operational excellence**, this is your opportunity to make a tangible impact in a high-stakes, fast-moving financial environment. Join us and help shape the future of resilient, efficient infrastructure.
يمكن أن يرتكب الذكاء الاصطناعي أخطاءً.
ℹ️ إخلاء مسؤولية توظيف:
موقع وظائف السعودية (ksajobshub.com) هو محرك بحث ومجمع لإعلانات الوظائف من المصادر والشركات الرسمية في المملكة العربية السعودية. نحن لا نتقاضى أي مبالغ مالية أو رسوم من الباحثين عن عمل، وتتم عمليات التقديم مباشرة عبر الانتقال للرابط الأصلي للجهة المعلنة.