HPC Senior Systems Administrator

KAUST (King Abdullah University of Science and Technology)
مكة المكرمة, مكة المكرمة دوام كامل
نشر: 1448/3/21 | 2026/09/03 ينتهي: 1448/4/22 | 2026/10/03 ✨ وصف بالذكاء الاصطناعي
تقدم للوظيفة الآن مشاركة عبر واتساب

الوصف الوظيفي

HPC Senior Systems Administrator – Lead High-Performance Computing Infrastructure at KAUST

Join the KAUST Supercomputing Laboratory (KSL) as a **Senior HPC Systems Administrator**, where you will play a pivotal role in maintaining, optimizing, and innovating one of the most advanced high-performance computing (HPC) environments in the region. This position offers an exceptional opportunity to work at the intersection of cutting-edge technology, computational research, and collaborative scientific discovery. As a key member of the KSL team, you will oversee the seamless operation of a sophisticated HPC infrastructure, ensuring world-class performance for researchers, engineers, and data scientists across computational science, engineering, big data analytics, and artificial intelligence/machine learning (AI/ML).

In this dynamic role, you will be responsible for the end-to-end management of a robust HPC ecosystem, including 600+ CPU and GPU nodes, high-performance storage systems, high-speed InfiniBand and Ethernet networks, and critical operational services. Your expertise will directly empower groundbreaking research initiatives, fostering innovation in fields such as computational fluid dynamics, materials science, climate modeling, and AI-driven simulations. This is more than a technical role—it’s a chance to be a strategic partner in advancing scientific and industrial breakthroughs.

Key Responsibilities

Your work will encompass a broad spectrum of technical and operational duties, ensuring the reliability, security, and high performance of KSL’s infrastructure. Key responsibilities include:

  • User Support and Service Excellence: Deliver exceptional support to researchers, faculty, and industry partners through multiple channels, including email, ticketing systems, and direct consultations. Maintain high standards of responsiveness, professionalism, and technical expertise to resolve complex issues efficiently.
  • Infrastructure Management: Install, configure, and administer HPC subsystems, including compute nodes, parallel file systems (e.g., Lustre, GPFS), InfiniBand and Ethernet networks, and configuration management tools (e.g., Ansible, Puppet). Ensure seamless integration and optimization of these components to support diverse workloads.
  • Cluster Operations and Automation: Deploy and manage cluster management software, monitoring tools, and workload schedulers such as Slurm, enforcing quality-of-service (QoS) policies, account management, and automation scripts (Python and C++). Develop and refine Bash and Python scripts to automate administrative tasks, reducing operational overhead and improving system reliability.
  • Containerization and Workload Support: Implement and maintain container environments (Singularity/Apptainer, Docker) tailored for HPC workloads, enabling flexible and reproducible research environments. Collaborate with application support teams to ensure seamless integration of containerized tools and frameworks.
  • Performance Benchmarking and Optimization: Conduct regular benchmarking of HPC components, including CPUs, GPUs, memory, InfiniBand, and storage systems. Analyze performance data to identify tuning opportunities, optimize hardware and software configurations, and advocate for system enhancements that align with evolving research needs.
  • Security and Compliance: Enforce stringent security best practices, including node hardening, kernel patching, and compliance with industry standards. Protect sensitive research data and ensure the integrity of the HPC environment through proactive security measures.
  • Research Collaboration and Tool Development: Actively support research activities by developing custom software tools, utilities, and utilities to address specific computational challenges. Work closely with faculty, researchers, and industry partners to translate technical requirements into scalable solutions that accelerate discovery.
  • Technology Evaluation and Innovation: Lead proof-of-concept projects and technology evaluations, assessing emerging HPC advancements and industry best practices. Advocate for system enhancements that align with KSL’s strategic goals and drive continuous improvement in infrastructure capabilities.
  • Vendor and Stakeholder Coordination: Collaborate with vendors and third-party service providers to troubleshoot issues, resolve technical challenges, and ensure timely resolution of hardware and software-related concerns. Act as a liaison between technical teams, researchers, and external partners to facilitate seamless operations.
  • Documentation and Knowledge Sharing: Develop and maintain comprehensive user documentation, standard operating procedures (SOPs), and training materials in KSL’s internal wiki. Foster a culture of knowledge sharing by documenting best practices, troubleshooting guides, and system configurations for current and future team members.
  • Strategic Leadership and Continuous Learning: Stay at the forefront of HPC advancements by engaging in continuous learning, attending industry conferences, and participating in professional networks. Drive benchmarking initiatives and contribute to future hardware procurement decisions based on data-driven insights and emerging trends.

Why This Role Matters

As a Senior HPC Systems Administrator at KSL, you will be at the heart of transformative research, enabling scientists and engineers to push the boundaries of what’s possible. Your technical expertise will directly impact the success of high-impact projects, from climate modeling and drug discovery to AI-driven simulations and materials science. This is an opportunity to make a tangible difference in advancing knowledge while working in a collaborative, interdisciplinary environment that values innovation and excellence.

Qualifications and Competencies

Candidates for this position should possess a robust blend of technical skills, research acumen, and a passion for HPC. Essential qualifications include:

  • Expertise in HPC User Support: Proven ability to provide technical support to users in computational science, engineering, data analysis, and AI/ML domains across diverse HPC environments.
  • Advanced Linux System Administration: Deep expertise in managing large-scale HPC environments using RHEL, Rocky Linux, or CentOS, with a focus on scalability, reliability, and performance optimization.
  • HPC Applications and Programming: Proficiency in HPC applications and programming models, including Fortran, C/C++, Python, MPI, OpenMP, CUDA, and OpenACC. Familiarity with high-performance computing libraries and frameworks used in research environments.
  • Complex HPC System Management: Hands-on experience managing intricate HPC systems, including parallel file systems (e.g., Lustre, GPFS), job schedulers (e.g., Slurm), InfiniBand/Ethernet networks, and monitoring systems (e.g., Ganglia, Prometheus).
  • Configuration Management: Experience with configuration management tools such as Ansible, Puppet, or Chef, enabling automated deployment, scaling, and maintenance of HPC infrastructure.
  • Research Collaboration: A track record of supporting research activities in collaborative HPC environments, with a strong ability to translate technical requirements into actionable solutions.
  • Analytical and Problem-Solving Skills: Strong analytical mindset with a proven ability to diagnose complex technical issues, implement solutions, and drive system improvements. Initiative and ownership are key traits for this role.
  • Project and Stakeholder Management: Experience with project management principles, including prioritization, deadline adherence, and cross-functional collaboration with researchers, vendors, and technical teams.
  • Innovation and Continuous Improvement: A proactive approach to identifying opportunities for system enhancements, benchmarking initiatives, and staying ahead of industry trends to ensure KSL remains at the forefront of HPC innovation.

If you are a seasoned HPC professional with a passion for enabling scientific discovery through robust infrastructure management, we invite you to apply and be part of KSL’s mission to advance knowledge through high-performance computing.

يمكن أن يرتكب الذكاء الاصطناعي أخطاءً.

المصدر: لينكد إن ↗ • 23 مشاهدة

ℹ️ إخلاء مسؤولية توظيف:

موقع وظائف السعودية (ksajobshub.com) هو محرك بحث ومجمع لإعلانات الوظائف من المصادر والشركات الرسمية في المملكة العربية السعودية. نحن لا نتقاضى أي مبالغ مالية أو رسوم من الباحثين عن عمل، وتتم عمليات التقديم مباشرة عبر الانتقال للرابط الأصلي للجهة المعلنة.

وظائف مشابهة

تقدم للوظيفة الآن