Role summary
The Data Center Hardware & Equipment Specialist at KAUST's Scientific Computing Center plays a crucial role in ensuring the physical infrastructure's security and operational continuity. This position oversees the installation, configuration, and maintenance of servers, including GPUs, within the CRAY supercomputer environment and other hosted equipment. The role ensures compliance with hardware tracking processes and updates the asset register while collaborating with various teams to meet procurement and regulatory requirements while maintaining high standards of security and operational efficiency.
Responsibilities
- Oversee the installation, configuration, and maintenance of servers, including GPUs
- Ensure compliance with hardware tracking processes and update the asset register
- Perform regular inspections and maintenance of data center hardware
- Monitor physical conditions of servers and other IT infrastructure
- Manage network cables, server hardware and other equipment
- Collaborate with vendors to proactively manage backup part inventory to ensure uptime
- Work with compliance officers to ensure accurate accounting of controlled equipment
- Follow data center access protocols to ensure the security of the system and prevent unauthorized removal of parts
- Respond to alarms and incidents, providing immediate resolution
- Collaborate with procurement teams to acquire necessary equipment
- Ensure adherence to export control regulations and NIST SP 800-53 standards
- Maintain detailed records of hardware assets and track equipment status
- Oversee or perform complex software and hardware troubleshooting, patches, and reinstallations
- Manage infrastructure capacity and performance, verifying application logs and monitoring activity
Qualifications
- Bachelor's degree in Computer Science, Information Technology, Electrical Engineering, Electronics Engineering or related field
- Relevant certifications preferred such as CDCTP, DCCA, CCNP Data Center, or RCDD
Requirements
- Detailed knowledge of HPE and CRAY supercomputer hardware infrastructure and NVIDIA supercomputer GPUs, including installation, maintenance, and troubleshooting of the supporting infrastructure
- Experience working in data centers, managing large-scale hardware deployments, and ensuring uptime and reliability
- Proven track record in overseeing the installation, configuration, and maintenance of servers and data center equipment
- Familiarity with hardware tracking processes and asset register management
- Minimum of 7 years of experience in managing HPC hardware and data center equipment and infrastructure
- Knowledge of data center infrastructure and operations
- Understanding of IT asset management for controlled equipment
Education
Bachelor's degree in Computer Science, Information Technology, Electrical Engineering, Electronics Engineering or related field
Experience
Minimum of 7 years of experience in managing HPC hardware and data center equipment and infrastructure
Skills
- Expertise in managing and maintaining HPE and CRAY supercomputer hardware infrastructure, including servers, storage systems, and networking equipment
- Proficiency in handling GPUs and optimizing infrastructure performance within the supercomputer environment
- Strong understanding of HPC systems and architectures
- Effective communication with stakeholders
- Use of monitoring tools to optimize supercomputer infrastructure performance
- Analytical skills to troubleshoot and resolve hardware issues
- Flexibility to adapt to new technologies and changing business needs
- Ensures compliance with hardware tracking processes
- Collaboration across teams to achieve goals
- Ability to manage logistics of heavy equipment and work in confined spaces