
Cluster Engineer
Posted Aug 6

Posted Aug 6
This is a fully remote position, open to applicants in United States.
β’ Design, implement, and enhance multi-node GPU clusters for AI training and inference tasks.
β’ Optimize distributed training setups to maximize GPU usage, throughput, and scaling efficiency.
β’ Improve inference clusters for token generation throughput, minimal latency, and high GPU utilization.
β’ Construct and maintain production AI infrastructure that operates hundreds to thousands of GPUs.
β’ Analyze and resolve performance bottlenecks across computing, networking, storage, and software layers.
β’ Conduct NCCL benchmarking, analysis, and tuning to enhance collective communication performance.
β’ Design and refine GPU networking utilizing InfiniBand or RoCE v2, focusing on RDMA, congestion management, topology awareness, and QoS.
β’ Configure and optimize distributed AI software stacks, including PyTorch, NCCL, CUDA, UCX, MPI, Slurm, and Pyxis/Enroot.
β’ Streamline GPU scheduling and resource allocation for both training and inference scenarios.
β’ Develop benchmarking and validation protocols for hardware, firmware, drivers, and software releases.
β’ Detect performance regressions and diagnose distributed training challenges at scale.
β’ Enhance storage architectures for checkpointing, dataset streaming, and high-performance parallel I/O.
β’ Collaborate closely with ML engineers to boost training scalability and inference effectiveness.
β’ Create automation for the deployment, validation, benchmarking, and monitoring of GPU clusters.
β’ Assess emerging AI infrastructure technologies and propose enhancements to platform architecture.
β’ Over 7 years of experience in designing or managing large-scale Linux infrastructure.
β’ More than 5 years of experience supporting production GPU clusters for AI or HPC workloads.
β’ Proven experience in building multi-node GPU training environments from scratch.
β’ Deep knowledge of distributed PyTorch training.
β’ Extensive experience in troubleshooting and optimizing NCCL communications.
β’ Understanding of AllReduce, ReduceScatter, AllGather, Broadcast, and point-to-point communications.
β’ Experience in benchmarking distributed training using nccl-tests, NVIDIA DCGM, and Nsight Systems; MLPerf is preferred.
β’ Knowledge of GPU memory management, KV Cache, activation checkpointing, tensor parallelism, pipeline parallelism, and data parallelism.
β’ Experience in optimizing LLM inference throughput, including tokens/sec, batch sizing, continuous batching, KV cache, and memory bandwidth.
β’ Proficiency in tuning CUDA, NCCL, UCX, and MPI.
β’ Expert-level skills in Linux systems administration.
β’ Familiarity with Slurm.
β’ Experience utilizing Pyxis and Enroot for containerized GPU workloads.
β’ Strong scripting skills in Python and Bash.
β’ Robust understanding of InfiniBand, RoCE v2, RDMA, GPUDirect RDMA, GPUDirect Storage, UCX, MPI, network topology optimization, congestion control, QoS, ECN/PFC, and high-speed Ethernet.
β’ Experience in designing or tuning AI storage solutions, including parallel file systems, distributed storage, object storage, NVMe, checkpoint optimization, dataset staging, storage bandwidth, and metadata performance.
β’ Expertise in NCCL benchmarking, multi-node scaling analysis, GPU utilization optimization, communication/computation overlap, NUMA optimization, CPU affinity, PCIe topology, GPU topology, memory bandwidth analysis, and end-to-end performance profiling.
β’ Preferred qualifications include experience with Kubernetes, NVIDIA GPU Operator, Kubernetes batch scheduling, vLLM/TensorRT-LLM/SGLang, NVIDIA DGX SuperPOD or similar, MLPerf, Prometheus/Grafana/DCGM Exporter, Ansible/Terraform, and AWS/Azure/GCP GPU environments.
β’ Competitive salary and performance-based bonuses.
β’ Comprehensive health, dental, and vision insurance.
β’ Opportunities for professional development and continuous learning.
β’ Flexible working hours and remote work options.
β’ Collaborative and innovative work environment.
CodiLime
Coderio
Get handpicked remote jobs straight to your inbox weekly.