Cluster Engineer

Posted Aug 6

This is a fully remote position, open to applicants in United States.

πŸ“‹ Description

β€’ Design, implement, and enhance multi-node GPU clusters for AI training and inference tasks.

β€’ Optimize distributed training setups to maximize GPU usage, throughput, and scaling efficiency.

β€’ Improve inference clusters for token generation throughput, minimal latency, and high GPU utilization.

β€’ Construct and maintain production AI infrastructure that operates hundreds to thousands of GPUs.

β€’ Analyze and resolve performance bottlenecks across computing, networking, storage, and software layers.

β€’ Conduct NCCL benchmarking, analysis, and tuning to enhance collective communication performance.

β€’ Design and refine GPU networking utilizing InfiniBand or RoCE v2, focusing on RDMA, congestion management, topology awareness, and QoS.

β€’ Configure and optimize distributed AI software stacks, including PyTorch, NCCL, CUDA, UCX, MPI, Slurm, and Pyxis/Enroot.

β€’ Streamline GPU scheduling and resource allocation for both training and inference scenarios.

β€’ Develop benchmarking and validation protocols for hardware, firmware, drivers, and software releases.

β€’ Detect performance regressions and diagnose distributed training challenges at scale.

β€’ Enhance storage architectures for checkpointing, dataset streaming, and high-performance parallel I/O.

β€’ Collaborate closely with ML engineers to boost training scalability and inference effectiveness.

β€’ Create automation for the deployment, validation, benchmarking, and monitoring of GPU clusters.

β€’ Assess emerging AI infrastructure technologies and propose enhancements to platform architecture.


⛳️ Requirements

β€’ Over 7 years of experience in designing or managing large-scale Linux infrastructure.

β€’ More than 5 years of experience supporting production GPU clusters for AI or HPC workloads.

β€’ Proven experience in building multi-node GPU training environments from scratch.

β€’ Deep knowledge of distributed PyTorch training.

β€’ Extensive experience in troubleshooting and optimizing NCCL communications.

β€’ Understanding of AllReduce, ReduceScatter, AllGather, Broadcast, and point-to-point communications.

β€’ Experience in benchmarking distributed training using nccl-tests, NVIDIA DCGM, and Nsight Systems; MLPerf is preferred.

β€’ Knowledge of GPU memory management, KV Cache, activation checkpointing, tensor parallelism, pipeline parallelism, and data parallelism.

β€’ Experience in optimizing LLM inference throughput, including tokens/sec, batch sizing, continuous batching, KV cache, and memory bandwidth.

β€’ Proficiency in tuning CUDA, NCCL, UCX, and MPI.

β€’ Expert-level skills in Linux systems administration.

β€’ Familiarity with Slurm.

β€’ Experience utilizing Pyxis and Enroot for containerized GPU workloads.

β€’ Strong scripting skills in Python and Bash.

β€’ Robust understanding of InfiniBand, RoCE v2, RDMA, GPUDirect RDMA, GPUDirect Storage, UCX, MPI, network topology optimization, congestion control, QoS, ECN/PFC, and high-speed Ethernet.

β€’ Experience in designing or tuning AI storage solutions, including parallel file systems, distributed storage, object storage, NVMe, checkpoint optimization, dataset staging, storage bandwidth, and metadata performance.

β€’ Expertise in NCCL benchmarking, multi-node scaling analysis, GPU utilization optimization, communication/computation overlap, NUMA optimization, CPU affinity, PCIe topology, GPU topology, memory bandwidth analysis, and end-to-end performance profiling.

β€’ Preferred qualifications include experience with Kubernetes, NVIDIA GPU Operator, Kubernetes batch scheduling, vLLM/TensorRT-LLM/SGLang, NVIDIA DGX SuperPOD or similar, MLPerf, Prometheus/Grafana/DCGM Exporter, Ansible/Terraform, and AWS/Azure/GCP GPU environments.


🏝️ Benefits

β€’ Competitive salary and performance-based bonuses.

β€’ Comprehensive health, dental, and vision insurance.

β€’ Opportunities for professional development and continuous learning.

β€’ Flexible working hours and remote work options.

β€’ Collaborative and innovative work environment.

People also viewed

CodiLime19 hours ago

Platform/System Validation Engineer, RDMA

PL flagPoland OnlyFreelanceEngineerPLN 22k – PLN 30k/month
ApplyView job
Coderio19 hours ago

Senior IA Engineer

AR flagArgentina, +4 more countriesFreelanceEngineer
ApplyView job
Interview Pen20 hours ago

Interview Engineer

AT flagAustria OnlyFreelanceEngineer
ApplyView job
Livefront20 hours ago

Agentic Engineer

PE flagPeru OnlyFull-timeEngineer
ApplyView job
Imubit20 hours ago

Refining and Petrochemical Process Optimization Engineer

US flagUnited States OnlyFull-timeEngineer
ApplyView job
CGWS - COME GROW WITH US20 hours ago

Senior Agentic Engineer

US flagUtah OnlyFull-timeEngineer
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers