
Senior HPC Cluster Engineer – AI, ML
Posted Aug 6

Posted Aug 6
This is a fully remote position, open to applicants in India.
• Lead in the administration of systems and service delivery within the AI/HPC fleet by managing system upgrades, addressing incidents, and enhancing reliability.
• Collaborate with international teams to provide an exceptional user experience in AI and HPC research.
• Oversee daily operations of production AI/HPC clusters, ensuring system performance, user satisfaction, and optimal resource utilization.
• Advance and enhance the ecosystem surrounding GPU-accelerated computing, including the development of scalable automation solutions.
• Construct and maintain diverse AI/ML clusters both on-premises and in cloud environments.
• Establish and nurture customer relationships and inter-team collaborations to adapt to evolving user requirements.
• Assist researchers in executing workloads, which includes performance assessment and optimization.
• Evaluate and enhance cluster efficiency, job fragmentation, and GPU resource usage to meet internal service level agreement (SLA) targets.
• Perform root cause analysis and recommend corrective measures.
• Take proactive measures to identify and resolve potential issues before they arise.
• Lead SEV triage and postmortem reviews for reliability incidents that impact users or infrastructure.
• Engage in on-call duty and incident response for critical production GPU clusters.
• Bachelor’s degree in Computer Science, Electrical Engineering, or a related field, or equivalent experience.
• At least 5 years of experience in designing and managing large-scale computing infrastructure.
• Familiarity with advanced AI/HPC job schedulers such as Slurm, K8s, PBS, RTDA, BCM, or LSF.
• Proficient in managing CentOS/RHEL and/or Ubuntu Linux distributions.
• Strong knowledge of cluster configuration management tools, including BCM, Terraform, Ansible, Puppet, and Salt.
• Understanding of container technologies such as Docker, Singularity, Podman, Shifter, and Charliecloud.
• Experience in Python programming and Bash scripting.
• Practical experience with AI/HPC workflows utilizing MPI.
• Experience in analyzing and tuning performance for diverse AI/HPC workloads.
• A commitment to continuous learning and staying informed about new HPC and AI/ML infrastructure technologies and methodologies.
• Familiarity with NVIDIA GPUs, CUDA programming, NCCL, and MLPerf benchmarking.
• Knowledge of AI/ML concepts, algorithms, models, and frameworks like PyTorch and TensorFlow.
• Proficient in InfiniBand, IPoIB, and RDMA technologies.
• Understanding of fast, distributed storage systems such as Lustre and GPFS for AI/HPC applications.
• Competitive salary and performance-based bonuses.
• Comprehensive health, dental, and vision insurance.
• Opportunities for professional development and training.
• Flexible working hours and potential remote work options.
• Access to cutting-edge technology and resources.
Blend360
Get handpicked remote jobs straight to your inbox weekly.