HPC Senior Systems Administrator
الوصف الوظيفي
HPC Senior Systems Administrator - KAUST Supercomputing Laboratory
The KAUST Supercomputing Laboratory (KSL) is seeking a highly motivated and skilled Senior HPC Systems Administrator to manage and support our state-of-the-art HPC cluster. This role offers a unique opportunity to provide broad support to researchers and end-users across computational science, engineering, big data analysis, and artificial intelligence/machine learning workloads.
- Responsibilities:
- Provide timely and effective user support via multiple channels, maintaining high customer service standards.
- Install, configure, and manage HPC subsystems, including compute nodes, storage systems, networks, and configuration management tools.
- Deploy and manage cluster management software, monitoring tools, and supporting services.
- Install and administer the Slurm workload manager, manage QOS policies, accounts, and related automation scripts.
- Develop and maintain automation scripts in Bash and Python to streamline system administration tasks.
- Deploy and manage container environments for HPC workloads.
- Benchmark HPC system components to ensure optimal performance and identify tuning opportunities.
- Enforce security best practices across all systems.
- Manage parallel file systems, including performance tuning and capacity planning.
- Support research activities by working closely with faculty, researchers, and industrial partners.
- Develop software tools and utilities to support research projects.
- Drive proof-of-concept projects and technology evaluations, and research industry best practices.
- Coordinate with vendors and third-party service providers to resolve issues.
- Develop and maintain user documentation, standard operating procedures, and training materials.
- Stay at the forefront of HPC advancements through continuous learning and collaboration.
Competencies: Expertise in supporting users of computational science and engineering, data analysis, and artificial intelligence applications in different HPC environments. Strong expertise in Linux system administration in large-scale HPC environments. Proficiency with HPC applications and programming models. Demonstrated track record of managing complex HPC systems. Experience with configuration management tools. Familiarity with computational science, data analysis, and AI/ML applications and libraries. Knowledge of project management principles. Demonstrated ability to support research activities in a highly collaborative HPC environment. Strong analytical, problem-solving, and decision-making skills. Proactive approach to system improvements. Ability to manage multiple concurrent projects and deliver high-quality results within deadlines. Proven ability to collaborate cross-functionally with researchers and other stakeholders.
يمكن أن يرتكب الذكاء الاصطناعي أخطاءً.