
Cluster Engineer
Posted Aug 6

Posted Aug 6
This is a fully remote position, open to applicants in United States.
• Design, implement, and enhance multi-node GPU clusters for AI training and inference tasks.
• Optimize distributed training setups to maximize GPU usage, throughput, and scaling efficiency.
• Improve inference clusters for token generation throughput, minimal latency, and high GPU utilization.
• Construct and maintain production AI infrastructure that operates hundreds to thousands of GPUs.
• Analyze and resolve performance bottlenecks across computing, networking, storage, and software layers.
• Conduct NCCL benchmarking, analysis, and tuning to enhance collective communication performance.
• Design and refine GPU networking utilizing InfiniBand or RoCE v2, focusing on RDMA, congestion management, topology awareness, and QoS.
• Configure and optimize distributed AI software stacks, including PyTorch, NCCL, CUDA, UCX, MPI, Slurm, and Pyxis/Enroot.
• Streamline GPU scheduling and resource allocation for both training and inference scenarios.
• Develop benchmarking and validation protocols for hardware, firmware, drivers, and software releases.
• Detect performance regressions and diagnose distributed training challenges at scale.
• Enhance storage architectures for checkpointing, dataset streaming, and high-performance parallel I/O.
• Collaborate closely with ML engineers to boost training scalability and inference effectiveness.
• Create automation for the deployment, validation, benchmarking, and monitoring of GPU clusters.
• Assess emerging AI infrastructure technologies and propose enhancements to platform architecture.
• Over 7 years of experience in designing or managing large-scale Linux infrastructure.
• More than 5 years of experience supporting production GPU clusters for AI or HPC workloads.
• Proven experience in building multi-node GPU training environments from scratch.
• Deep knowledge of distributed PyTorch training.
• Extensive experience in troubleshooting and optimizing NCCL communications.
• Understanding of AllReduce, ReduceScatter, AllGather, Broadcast, and point-to-point communications.
• Experience in benchmarking distributed training using nccl-tests, NVIDIA DCGM, and Nsight Systems; MLPerf is preferred.
• Knowledge of GPU memory management, KV Cache, activation checkpointing, tensor parallelism, pipeline parallelism, and data parallelism.
• Experience in optimizing LLM inference throughput, including tokens/sec, batch sizing, continuous batching, KV cache, and memory bandwidth.
• Proficiency in tuning CUDA, NCCL, UCX, and MPI.
• Expert-level skills in Linux systems administration.
• Familiarity with Slurm.
• Experience utilizing Pyxis and Enroot for containerized GPU workloads.
• Strong scripting skills in Python and Bash.
• Robust understanding of InfiniBand, RoCE v2, RDMA, GPUDirect RDMA, GPUDirect Storage, UCX, MPI, network topology optimization, congestion control, QoS, ECN/PFC, and high-speed Ethernet.
• Experience in designing or tuning AI storage solutions, including parallel file systems, distributed storage, object storage, NVMe, checkpoint optimization, dataset staging, storage bandwidth, and metadata performance.
• Expertise in NCCL benchmarking, multi-node scaling analysis, GPU utilization optimization, communication/computation overlap, NUMA optimization, CPU affinity, PCIe topology, GPU topology, memory bandwidth analysis, and end-to-end performance profiling.
• Preferred qualifications include experience with Kubernetes, NVIDIA GPU Operator, Kubernetes batch scheduling, vLLM/TensorRT-LLM/SGLang, NVIDIA DGX SuperPOD or similar, MLPerf, Prometheus/Grafana/DCGM Exporter, Ansible/Terraform, and AWS/Azure/GCP GPU environments.
• Competitive salary and performance-based bonuses.
• Comprehensive health, dental, and vision insurance.
• Opportunities for professional development and continuous learning.
• Flexible working hours and remote work options.
• Collaborative and innovative work environment.
Highland Electric Fleets
Falconwood, Incorporated
Aira
Pragmatike
Get handpicked remote jobs straight to your inbox weekly.