Remotery

Cluster Engineer

Posted Aug 6

This is a fully remote position, open to applicants in United States.

📋 Description

• Design, implement, and enhance multi-node GPU clusters for AI training and inference tasks.

• Optimize distributed training setups to maximize GPU usage, throughput, and scaling efficiency.

• Improve inference clusters for token generation throughput, minimal latency, and high GPU utilization.

• Construct and maintain production AI infrastructure that operates hundreds to thousands of GPUs.

• Analyze and resolve performance bottlenecks across computing, networking, storage, and software layers.

• Conduct NCCL benchmarking, analysis, and tuning to enhance collective communication performance.

• Design and refine GPU networking utilizing InfiniBand or RoCE v2, focusing on RDMA, congestion management, topology awareness, and QoS.

• Configure and optimize distributed AI software stacks, including PyTorch, NCCL, CUDA, UCX, MPI, Slurm, and Pyxis/Enroot.

• Streamline GPU scheduling and resource allocation for both training and inference scenarios.

• Develop benchmarking and validation protocols for hardware, firmware, drivers, and software releases.

• Detect performance regressions and diagnose distributed training challenges at scale.

• Enhance storage architectures for checkpointing, dataset streaming, and high-performance parallel I/O.

• Collaborate closely with ML engineers to boost training scalability and inference effectiveness.

• Create automation for the deployment, validation, benchmarking, and monitoring of GPU clusters.

• Assess emerging AI infrastructure technologies and propose enhancements to platform architecture.


⛳️ Requirements

• Over 7 years of experience in designing or managing large-scale Linux infrastructure.

• More than 5 years of experience supporting production GPU clusters for AI or HPC workloads.

• Proven experience in building multi-node GPU training environments from scratch.

• Deep knowledge of distributed PyTorch training.

• Extensive experience in troubleshooting and optimizing NCCL communications.

• Understanding of AllReduce, ReduceScatter, AllGather, Broadcast, and point-to-point communications.

• Experience in benchmarking distributed training using nccl-tests, NVIDIA DCGM, and Nsight Systems; MLPerf is preferred.

• Knowledge of GPU memory management, KV Cache, activation checkpointing, tensor parallelism, pipeline parallelism, and data parallelism.

• Experience in optimizing LLM inference throughput, including tokens/sec, batch sizing, continuous batching, KV cache, and memory bandwidth.

• Proficiency in tuning CUDA, NCCL, UCX, and MPI.

• Expert-level skills in Linux systems administration.

• Familiarity with Slurm.

• Experience utilizing Pyxis and Enroot for containerized GPU workloads.

• Strong scripting skills in Python and Bash.

• Robust understanding of InfiniBand, RoCE v2, RDMA, GPUDirect RDMA, GPUDirect Storage, UCX, MPI, network topology optimization, congestion control, QoS, ECN/PFC, and high-speed Ethernet.

• Experience in designing or tuning AI storage solutions, including parallel file systems, distributed storage, object storage, NVMe, checkpoint optimization, dataset staging, storage bandwidth, and metadata performance.

• Expertise in NCCL benchmarking, multi-node scaling analysis, GPU utilization optimization, communication/computation overlap, NUMA optimization, CPU affinity, PCIe topology, GPU topology, memory bandwidth analysis, and end-to-end performance profiling.

• Preferred qualifications include experience with Kubernetes, NVIDIA GPU Operator, Kubernetes batch scheduling, vLLM/TensorRT-LLM/SGLang, NVIDIA DGX SuperPOD or similar, MLPerf, Prometheus/Grafana/DCGM Exporter, Ansible/Terraform, and AWS/Azure/GCP GPU environments.


🏝️ Benefits

• Competitive salary and performance-based bonuses.

• Comprehensive health, dental, and vision insurance.

• Opportunities for professional development and continuous learning.

• Flexible working hours and remote work options.

• Collaborative and innovative work environment.

People also viewed

Highland Electric Fleets12 hours ago

Energy Storage Project Engineer

US flagMassachusetts OnlyFull-timeEngineer$90k – $105k/year
ApplyView job
Falconwood, Incorporated12 hours ago

Endpoint Engineer

US flagUnited States OnlyFull-timeEngineer$130k – $140k/year
ApplyView job
Aira12 hours ago

Plumbing and Heating Engineer

GB flagUnited Kingdom OnlyFull-timeEngineer£36.1k/year
ApplyView job
Pragmatike13 hours ago

Forward Deployed Engineer, YC Experience

US flagCalifornia, +3 more statesFull-timeEngineer$200k – $350k/year
ApplyView job
Pragmatike13 hours ago

Lead Forward Deployed Engineer

US flagCalifornia, +2 more statesFull-timeEngineer$200k – $350k/year
ApplyView job
Pragmatike13 hours ago

Senior Forward Deployed Engineer

US flagCalifornia, +3 more statesFull-timeEngineer$200k – $350k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers