
Senior/Staff Kubernetes Infrastructure Engineer
Posted 19 hours ago

Posted 19 hours ago
This is a fully remote position, open to applicants in United States.
• Design, automate, validate, and oversee the entire lifecycle of customer compute environments, including provisioning, upgrades, recovery, and decommissioning.
• Leverage AI to streamline and expedite infrastructure delivery and operations.
• Provision dedicated Kubernetes and Slurm clusters customized to meet customer workloads.
• Create and maintain Linux images along with automated operating system provisioning workflows.
• Manage the NVIDIA GPU stack, which includes drivers, GPU Operator, NVIDIA Container Toolkit, device plugins, MIG, and GPU monitoring.
• Design Kubernetes and data center networking utilizing Cilium/Calico, MetalLB, VLAN, VXLAN, BGP, and ECMP.
• Configure distributed and shared storage solutions for high-performance workloads.
• Establish monitoring, alerting, diagnostics, and automated recovery systems for customer environments.
• Develop reusable tools, standards, documentation, and runbooks.
• Collaborate with customers and internal teams to convert workload requirements into effective infrastructure designs.
• Over 5 years of experience in building and managing production Linux infrastructure.
• Extensive production experience with Kubernetes on bare metal, including bootstrapping, upgrades, high availability control planes, etcd, containerd, CNI, CSI, ingress, load balancing, observability, security, and troubleshooting.
• Familiarity with Linux virtualization technologies such as KVM/QEMU, libvirt, and VFIO device passthrough.
• Experience in operating NVIDIA GPUs on Linux and Kubernetes, encompassing drivers, container runtimes, device plugins, GPU Operator, and GPU telemetry.
• Strong foundational knowledge of networking principles: TCP/IP, L2/L3, VLANs, routing, and packet-level troubleshooting using tcpdump and Wireshark.
• Practical experience in scripting.
• Proficiency with configuration management tools like Ansible.
• Capable of diagnosing intricate, cross-layer infrastructure issues.
• Excellent communication skills and the ability to influence technical decisions across teams.
• Proven history of acting swiftly, taking ownership, and continuously enhancing systems.
• Legally authorized to work in the United States.
• Nice-to-have: Experience with production Slurm.
• Nice-to-have: Experience in high-performance networking with NVLink/NVSwitch, InfiniBand, RoCEv2, GPUDirect RDMA, NCCL, or IMEX.
• Nice-to-have: Knowledge in Hugepages, NUMA, CPU pinning, SR-IOV, DPDK, Ceph, Lustre, Weka, KubeVirt, OpenStack, IPsec, WireGuard, Tailscale, VXLAN, BGP, ECMP, BMC, IPMI, Redfish, PXE/iPXE, Kickstart, cloud-init, NetBox, Nautobot, Nornir, AI training/inference/distributed GPU workload infrastructure, or proficiency in Python/Go.
• Equity
• Salary range of $180K–$250K
Trilon Group
Carbon60
LatamCent
AIP Publishing
Get handpicked remote jobs straight to your inbox weekly.