
Senior GPU Cloud, K8S Expert
Posted Sep 11

Posted Sep 11
This is a fully remote position, open to applicants in California, +1 more state.
• Design, implement, and manage the control plane for an AI-driven GPU cloud.
• Operate production Kubernetes clusters tailored for GPU workloads at scale (ranging from 100 to 10,000 GPUs).
• Configure Nvidia GPU operator, device plugin, MIG, and policies for GPU time-slicing.
• Execute topology-aware scheduling that leverages GPU locality, NVLink domain awareness, and network rail affinity.
• Create Custom Resource Definitions (CRDs) to manage the lifecycle of GPU workloads.
• Integrate Slurm, Ray, and Kubeflow with Kubernetes.
• Ensure multi-tenant isolation through the use of namespaces, network policies, resource quotas, RBAC, and pod security standards.
• Automate the provisioning of BMaaS, tenant onboarding, lifecycle management, and reclamation processes.
• Develop Terraform providers and modules to facilitate infrastructure-as-code across GPU clusters.
• Establish and achieve SLIs/SLOs concerning cluster availability, job completion rates, and provisioning latency.
• Streamline incident management, escalation protocols, and post-incident review processes.
• Oversee monitoring systems using Prometheus, Grafana, Alertmanager, and PagerDuty.
• Automate the detection of GPU node failures, including draining, cordoning, tainting, and workload rescheduling.
• Ensure the control plane is secure for automated AIOps remediation.
• Facilitate automated workflows for anticipated GPU faults without affecting customers.
• Launch BMaaS for external clients, enabling self-service onboarding.
• Publish and ensure adherence to cluster availability and job-completion SLOs.
• A minimum of 5 years in Kubernetes operations, including at least 2 years focused on managing GPU workloads within K8S.
• Profound knowledge of Nvidia GPU operator, device plugin, and GPU scheduling in Kubernetes.
• Experience with topology-aware scheduling and management of GPU-specific resources.
• Practical experience in building multi-tenant Kubernetes platforms with robust isolation guarantees.
• Background in bare-metal server provisioning and lifecycle automation (Ironic, MAAS, or custom solutions).
• Expertise in Terraform, Helm, and GitOps workflows (ArgoCD/Flux).
• A solid SRE foundation, including familiarity with SLI/SLO frameworks, incident management, and capacity planning.
• Experience working with Prometheus, Grafana, and large-scale alerting systems.
• Strong programming skills in Go or Python for the development of operators/CRDs.
• An aptitude for AIOps — viewing the Kubernetes control plane as a platform for automated remediation rather than just a scheduler.
• Experience in integrating an autoscaler/remediator loop into Kubernetes or the capability to design one.
• A runbook-as-code mentality — every SRE playbook you create should be executable by the platform.
• Competitive salary and performance-based bonuses.
• Comprehensive health, dental, and vision insurance.
• Flexible work hours and remote work options.
• Opportunities for professional development and continuous learning.
• A collaborative and inclusive work environment.
FourEnergy GmbH
ICF
Mastercam
C&S Informática
Get handpicked remote jobs straight to your inbox weekly.