Remotery

Senior/Staff Kubernetes Infrastructure Engineer

atfalRemoteUS flagUnited StatesFull-timeInfrastructure EngineerSenior$180k – $250k/year

Posted 19 hours ago

This is a fully remote position, open to applicants in United States.

📋 Description

• Design, automate, validate, and oversee the entire lifecycle of customer compute environments, including provisioning, upgrades, recovery, and decommissioning.

• Leverage AI to streamline and expedite infrastructure delivery and operations.

• Provision dedicated Kubernetes and Slurm clusters customized to meet customer workloads.

• Create and maintain Linux images along with automated operating system provisioning workflows.

• Manage the NVIDIA GPU stack, which includes drivers, GPU Operator, NVIDIA Container Toolkit, device plugins, MIG, and GPU monitoring.

• Design Kubernetes and data center networking utilizing Cilium/Calico, MetalLB, VLAN, VXLAN, BGP, and ECMP.

• Configure distributed and shared storage solutions for high-performance workloads.

• Establish monitoring, alerting, diagnostics, and automated recovery systems for customer environments.

• Develop reusable tools, standards, documentation, and runbooks.

• Collaborate with customers and internal teams to convert workload requirements into effective infrastructure designs.


⛳️ Requirements

• Over 5 years of experience in building and managing production Linux infrastructure.

• Extensive production experience with Kubernetes on bare metal, including bootstrapping, upgrades, high availability control planes, etcd, containerd, CNI, CSI, ingress, load balancing, observability, security, and troubleshooting.

• Familiarity with Linux virtualization technologies such as KVM/QEMU, libvirt, and VFIO device passthrough.

• Experience in operating NVIDIA GPUs on Linux and Kubernetes, encompassing drivers, container runtimes, device plugins, GPU Operator, and GPU telemetry.

• Strong foundational knowledge of networking principles: TCP/IP, L2/L3, VLANs, routing, and packet-level troubleshooting using tcpdump and Wireshark.

• Practical experience in scripting.

• Proficiency with configuration management tools like Ansible.

• Capable of diagnosing intricate, cross-layer infrastructure issues.

• Excellent communication skills and the ability to influence technical decisions across teams.

• Proven history of acting swiftly, taking ownership, and continuously enhancing systems.

• Legally authorized to work in the United States.

• Nice-to-have: Experience with production Slurm.

• Nice-to-have: Experience in high-performance networking with NVLink/NVSwitch, InfiniBand, RoCEv2, GPUDirect RDMA, NCCL, or IMEX.

• Nice-to-have: Knowledge in Hugepages, NUMA, CPU pinning, SR-IOV, DPDK, Ceph, Lustre, Weka, KubeVirt, OpenStack, IPsec, WireGuard, Tailscale, VXLAN, BGP, ECMP, BMC, IPMI, Redfish, PXE/iPXE, Kickstart, cloud-init, NetBox, Nautobot, Nornir, AI training/inference/distributed GPU workload infrastructure, or proficiency in Python/Go.


🏝️ Benefits

• Equity

• Salary range of $180K–$250K

People also viewed

Trilon Group14 hours ago

AWS Infrastructure Engineer

US flagUnited States OnlyFull-timeInfrastructure Engineer$90k – $110k/year
ApplyView job
Carbon6018 hours ago

Principal Infrastructure Architect

CA flagCanada OnlyFull-timeInfrastructure EngineerC$180k – C$220k/year
ApplyView job
LatamCent23 hours ago

Lead Security and Infrastructure Engineer

US flagFlorida OnlyFull-timeInfrastructure Engineer$170k – $210k/year
ApplyView job
AIP Publishing1 day ago

Cloud Infrastructure Engineer

US flagConnecticut, +9 more statesFull-timeInfrastructure Engineer$125k – $135k/year
ApplyView job
AECOM1 day ago

Development Infrastructure Engineer

GB flagUnited Kingdom OnlyFull-timeInfrastructure Engineer
ApplyView job
FCamara Consulting & Training1 day ago

Senior Infrastructure Engineer

PT flagPortugal OnlyFull-timeInfrastructure Engineer
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers