Role summary
A key member of KAUST's Supercomputing Lab AI/ML Support Team, supporting the delivery of AI research services to KAUST's diverse research community. Working under the AI/ML Support Team Lead, this role focuses on developing and optimizing Generative AI models, maintaining computational benchmarks, and providing expert consultation to researchers across multiple scientific domains, including Climate & Weather, Bioinformatics, CFD, NLP, and multimodal AI.
Responsibilities
- Providing timely and useful user support via telephone, walk-in, email, and ticketing system submissions for all types of inquiries
- Maintaining high customer service standards in dealing with and responding to user issues and questions
- Developing and consulting on large-scale Generative AI model training on domain-specific datasets across research areas including Climate & Weather, Bioinformatics, CFD, NLP, and multimodal AI
- Supporting researchers in fine-tuning foundation models on domain-specific datasets using advanced optimization techniques
- Developing data engineering pipelines to support AI research workflows
- Designing and implementing efficient AI workflows optimized for KSL's high-performance computing environment
- Building and maintaining secure, OCI-compliant, HPC-ready container images using Singularity, Podman, or similar
- Developing complex workflows using SLURM and Kubernetes for distributed training and inference
- Conducting computational readiness reviews for AI research projects
- Assisting in AI model and artifact control reviews to ensure compliance with institutional standards
- Supporting researchers in designing secure, compliant, and performant workflows
- Providing expert consultation to researchers on efficient utilization of AI resources and best practices
- Supporting the implementation of usage monitoring and reporting systems
- Ensuring user workflows comply with KSL security policies and best practices
- Developing and maintaining computational benchmarks for AI workloads on KSL systems
- Creating and maintaining regression testing workloads to stress test system functionality
- Supporting performance debugging and optimization activities for research workloads
- Contributing to technology evaluation and benchmarking exercises for future infrastructure investments
- Performing benchmarking of new hardware and software configurations
- Creating comprehensive training materials for end-users on KSL's HPC systems hosting AI workloads and tools
- Developing and maintaining high-quality technical documentation
- Supporting the delivery of workshops on distributed training, fine-tuning, and inference optimization
- Contributing to knowledge transfer initiatives within the KAUST research community
- Providing one-on-one consultation to researchers on efficient use of computational resources
Qualifications
- Bachelor's or Master's degree in Computer Science, Data Science, Computational Science, Artificial Intelligence, or related field
- Strong academic foundation in machine learning, deep learning, and AI fundamentals
Education
Bachelor's or Master's degree in Computer Science, Data Science, Computational Science, Artificial Intelligence, or related field
Skills
- Programming: Proficiency in Python; experience with R, Julia, Rust or C/C++ is a plus
- AI/ML Frameworks: Strong expertise in PyTorch and/or TensorFlow, JAX or similar
- Generative AI: Experience with foundation model development and fine-tuning techniques
- HPC Systems: Experience developing complex workflows using SLURM and/or Kubernetes
- Containerization: Experience building efficient HPC-ready container images using Singularity, Podman or similar
- Data Engineering: Experience with data engineering techniques for developing AI pipelines
- Linux: Strong Linux/Unix skills and bash scripting capabilities
- Experience with Cray EX supercomputers with NVIDIA GPUs
- Experience with Kubeflow pipelines and Kubeflow Training Operator
- Experience with distributed inference frameworks (NVIDIA Triton, NIM, SGLang, llama.cpp, llm-d, LLMcache)
- Knowledge of security vulnerability inspection in software libraries, AI models, datasets, and pipelines
- Experience with software supply chain tools (JFrog, Nexus, Trivy, Cloudsmith)
- Experience with data management on S3-compatible object storage at scale
- Experience with high-performance distributed filesystems (Lustre, Weka IO, VAST Data)
- Proficiency with NVIDIA Nsight and Compute for profiling AI workloads on GPUs
- Experience developing CI/CD pipelines using GitLab, Travis, CircleCI, or similar tools
- Experience with software build tools (autoconf, CMake, scons, SPACK, EasyBuild, Conda, Pip)
- Strong problem-solving and analytical abilities
- Excellent written and verbal communication skills in English
- Customer service mindset with patience for supporting diverse skill levels
- Ability to work independently and as part of a collaborative team
- Strong documentation and knowledge-sharing practices
- Cultural sensitivity for working in an international environment