Role summary
Lead a team ensuring smooth operation of a Linux cluster comprising 300+ GPU/CPU compute nodes with parallel filesystems and high-performance networks. This technical and people leadership role involves supervising 3-4 experienced HPC systems administrators and developing standard operating procedures.
Responsibilities
- Plan system operation and upgrades to meet laboratory and customer requirements
- Develop and implement workload scheduler policies
- Support high-performance filesystems
- Manage network infrastructure including TCP/IP and HPC networks
- Use scripting languages for node automation and configuration management
- Manage hardware failures and spare parts
- Build effective relationships with staff, faculty and students through Core Labs
- Manage multiple or significant projects using sophisticated planning techniques
- Identify technical training needs for attached staff
- Respond to security and safety incidents as a resource and team member
- Enhance technical methodology through expansion or development of new efforts
- Provide innovative problem-solving approaches to enhance organizational capabilities
- Understand and contribute to broad strategic objectives
- Nurture and maintain relationships with major customers
- Initiate new project concepts and develop technical proposals
- Supervise several scientists, engineers or technicians on assigned work
- Provide major input to project team staffing
- Build and optimize teams for efficiency and cost-effectiveness
- Identify and evaluate candidates for open positions
- Mentor and train staff in technical, project and business development skills
Qualifications
- Bachelor of Science or equivalent in relevant discipline plus 10 years' experience, OR Master of Science or equivalent in relevant discipline plus 7 years' experience, OR Doctor of Philosophy or equivalent in relevant discipline plus 5 years' experience
Skills
- SLURM workload manager including GPU scheduling
- Parallel filesystems (Weka IO, Lustre)
- TCP/IP and high performance networks (Infiniband)
- Proficiency in scripting languages (Bash, Python, Ruby)
- Familiarity with configuration management tools (Puppet)
- Expert documentation skills
- Working level contact with users and suppliers
- Analytical and systematic approach to problem solving
- Initiative in identifying and negotiating development opportunities
- Effective communication skills in written and oral English
- Ability to work effectively with other Supercomputing Laboratory teams
- Competent planning, scheduling and monitoring of own work and others' within deadlines
- Successful work in highly collaborative research environments
- Discretion in identifying and resolving complex problems
- Ability to perform broad range of complex and non-routine work in various environments
- Expert-level knowledge of laboratory systems
Education · Experience · Type
Bachelor'sRequires 5+ yearsSalary not stated