Remotery

Senior HPC Cluster Engineer – AI, ML

Posted Aug 6

This is a fully remote position, open to applicants in India.

📋 Description

• Lead in the administration of systems and service delivery within the AI/HPC fleet by managing system upgrades, addressing incidents, and enhancing reliability.

• Collaborate with international teams to provide an exceptional user experience in AI and HPC research.

• Oversee daily operations of production AI/HPC clusters, ensuring system performance, user satisfaction, and optimal resource utilization.

• Advance and enhance the ecosystem surrounding GPU-accelerated computing, including the development of scalable automation solutions.

• Construct and maintain diverse AI/ML clusters both on-premises and in cloud environments.

• Establish and nurture customer relationships and inter-team collaborations to adapt to evolving user requirements.

• Assist researchers in executing workloads, which includes performance assessment and optimization.

• Evaluate and enhance cluster efficiency, job fragmentation, and GPU resource usage to meet internal service level agreement (SLA) targets.

• Perform root cause analysis and recommend corrective measures.

• Take proactive measures to identify and resolve potential issues before they arise.

• Lead SEV triage and postmortem reviews for reliability incidents that impact users or infrastructure.

• Engage in on-call duty and incident response for critical production GPU clusters.


⛳️ Requirements

• Bachelor’s degree in Computer Science, Electrical Engineering, or a related field, or equivalent experience.

• At least 5 years of experience in designing and managing large-scale computing infrastructure.

• Familiarity with advanced AI/HPC job schedulers such as Slurm, K8s, PBS, RTDA, BCM, or LSF.

• Proficient in managing CentOS/RHEL and/or Ubuntu Linux distributions.

• Strong knowledge of cluster configuration management tools, including BCM, Terraform, Ansible, Puppet, and Salt.

• Understanding of container technologies such as Docker, Singularity, Podman, Shifter, and Charliecloud.

• Experience in Python programming and Bash scripting.

• Practical experience with AI/HPC workflows utilizing MPI.

• Experience in analyzing and tuning performance for diverse AI/HPC workloads.

• A commitment to continuous learning and staying informed about new HPC and AI/ML infrastructure technologies and methodologies.

• Familiarity with NVIDIA GPUs, CUDA programming, NCCL, and MLPerf benchmarking.

• Knowledge of AI/ML concepts, algorithms, models, and frameworks like PyTorch and TensorFlow.

• Proficient in InfiniBand, IPoIB, and RDMA technologies.

• Understanding of fast, distributed storage systems such as Lustre and GPFS for AI/HPC applications.


🏝️ Benefits

• Competitive salary and performance-based bonuses.

• Comprehensive health, dental, and vision insurance.

• Opportunities for professional development and training.

• Flexible working hours and potential remote work options.

• Access to cutting-edge technology and resources.

People also viewed

dexter health9 hours ago

Applied AI Engineer

DE flagGermany OnlyFull-timeAI Engineer
ApplyView job
Blend3609 hours ago

Lead AI Engineer – Agentic Engineering

IN flagIndia OnlyFull-timeAI Engineer
ApplyView job
CI&T9 hours ago

Senior Generative AI Developer

BR flagBrazil OnlyFull-timeAI Engineer
ApplyView job
Lingaro10 hours ago

ML/AI Engineer

PL flagPoland OnlyFreelanceAI Engineer
ApplyView job
Orion Innovation10 hours ago

Cloud/AI Developer

US flagUnited States OnlyFull-timeAI Engineer
ApplyView job
Tango11 hours ago

Senior Applied AI Engineer

US flagUnited States, +1 more stateFull-timeAI Engineer$160k – $190k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers