On-Premise DevOps Engineer
الوصف الوظيفي
On-Premise DevOps Engineer – Drive Innovation in Infrastructure and Automation
Are you a seasoned DevOps and Platform Engineer passionate about architecting, optimizing, and maintaining robust on-premise infrastructure? We are seeking a highly skilled On-Premise DevOps Engineer with 6+ years of hands-on experience to lead the design, deployment, and management of self-hosted CI/CD pipelines, containerized environments, and monitoring solutions. This role is ideal for an expert who thrives in a fast-paced, mission-critical environment where reliability, scalability, and automation are paramount.
In this strategic position, you will play a pivotal role in ensuring seamless operations across our on-premise infrastructure while driving continuous improvement in platform engineering practices. Your expertise will be instrumental in building resilient, high-performance systems that support our organization’s growth and innovation goals.
Key Responsibilities
As an On-Premise DevOps Engineer, your responsibilities will span a broad spectrum of technical disciplines, including:
- CI/CD Pipeline Architecture and Implementation: Design and deploy sophisticated CI/CD pipelines using Jenkins/CloudBees, integrating artifact management, automated code quality gates, and secure deployment workflows to ensure rapid, reliable, and auditable releases.
- Platform Engineering and Operations: Architect, administer, and maintain OpenShift clusters in production, focusing on capacity planning, upgrades, role-based access control (RBAC), networking, ingress management, and workload optimization to maximize performance and availability.
- Monitoring and Observability: Lead the implementation and optimization of monitoring solutions using Prometheus and Grafana, including advanced PromQL queries, custom dashboards, Service Level Objective (SLO) definitions, and proactive alert tuning to preemptively address infrastructure issues.
- Centralized Logging and Analytics: Configure and maintain ELK Stack (Elasticsearch, Logstash, Kibana) for centralized application logging, ensuring efficient log parsing, indexing, and retention policies to support troubleshooting and compliance requirements.
- Disaster Recovery and High Availability: Design and maintain robust PR-DR (Primary-Replica Disaster Recovery) capabilities, including replication strategies, failover/failback mechanisms, and periodic DR drills to minimize downtime and ensure business continuity.
- Database Administration: Manage critical databases such as PostgreSQL or MySQL, including installation, backup/restore procedures, replication setups, monitoring, and performance tuning to guarantee data integrity and operational efficiency.
- Backend Development and Integrations: Develop and maintain custom Python-based backend services and REST APIs, integrating them with existing systems to enhance automation, workflow efficiency, and platform capabilities.
- Infrastructure Optimization: Analyze resource utilization, rightsizing workloads, and providing data-driven recommendations to stakeholders to improve cost-effectiveness and performance across the infrastructure.
- Troubleshooting and Root Cause Analysis: Diagnose and resolve complex issues across the technology stack, including CI/CD failures, OpenShift cluster anomalies, storage and networking bottlenecks, and infrastructure-related challenges with minimal downtime.
- Documentation and Knowledge Sharing: Develop and maintain comprehensive technical runbooks, documentation, and best practices to foster a culture of transparency and continuous learning within the engineering team.
- Mentorship and Collaboration: Mentor junior engineers, share expertise, and contribute to evolving platform engineering practices to elevate the overall maturity and efficiency of the DevOps ecosystem.
Why This Role Matters
This position is not just about maintaining systems—it’s about shaping the future of our on-premise infrastructure. You will collaborate closely with cross-functional teams to align technology with business objectives, ensuring that our platforms are not only highly available but also scalable, secure, and future-proof. Your work will directly impact the reliability of our operations, enabling us to deliver value to our stakeholders with confidence.
If you are a detail-oriented, innovative thinker with a passion for automation, scalability, and system reliability, we invite you to join our team and make a meaningful impact in a dynamic and collaborative environment.
يمكن أن يرتكب الذكاء الاصطناعي أخطاءً.
ℹ️ إخلاء مسؤولية توظيف:
موقع وظائف السعودية (ksajobshub.com) هو محرك بحث ومجمع لإعلانات الوظائف من المصادر والشركات الرسمية في المملكة العربية السعودية. نحن لا نتقاضى أي مبالغ مالية أو رسوم من الباحثين عن عمل، وتتم عمليات التقديم مباشرة عبر الانتقال للرابط الأصلي للجهة المعلنة.