Senior GPU Cloud, K8S Expert

Posted Sep 11

This is a fully remote position, open to applicants in California, +1 more state.

📋 Description

• Design, implement, and manage the control plane for an AI-driven GPU cloud.

• Operate production Kubernetes clusters tailored for GPU workloads at scale (ranging from 100 to 10,000 GPUs).

• Configure Nvidia GPU operator, device plugin, MIG, and policies for GPU time-slicing.

• Execute topology-aware scheduling that leverages GPU locality, NVLink domain awareness, and network rail affinity.

• Create Custom Resource Definitions (CRDs) to manage the lifecycle of GPU workloads.

• Integrate Slurm, Ray, and Kubeflow with Kubernetes.

• Ensure multi-tenant isolation through the use of namespaces, network policies, resource quotas, RBAC, and pod security standards.

• Automate the provisioning of BMaaS, tenant onboarding, lifecycle management, and reclamation processes.

• Develop Terraform providers and modules to facilitate infrastructure-as-code across GPU clusters.

• Establish and achieve SLIs/SLOs concerning cluster availability, job completion rates, and provisioning latency.

• Streamline incident management, escalation protocols, and post-incident review processes.

• Oversee monitoring systems using Prometheus, Grafana, Alertmanager, and PagerDuty.

• Automate the detection of GPU node failures, including draining, cordoning, tainting, and workload rescheduling.

• Ensure the control plane is secure for automated AIOps remediation.

• Facilitate automated workflows for anticipated GPU faults without affecting customers.

• Launch BMaaS for external clients, enabling self-service onboarding.

• Publish and ensure adherence to cluster availability and job-completion SLOs.


⛳️ Requirements

• A minimum of 5 years in Kubernetes operations, including at least 2 years focused on managing GPU workloads within K8S.

• Profound knowledge of Nvidia GPU operator, device plugin, and GPU scheduling in Kubernetes.

• Experience with topology-aware scheduling and management of GPU-specific resources.

• Practical experience in building multi-tenant Kubernetes platforms with robust isolation guarantees.

• Background in bare-metal server provisioning and lifecycle automation (Ironic, MAAS, or custom solutions).

• Expertise in Terraform, Helm, and GitOps workflows (ArgoCD/Flux).

• A solid SRE foundation, including familiarity with SLI/SLO frameworks, incident management, and capacity planning.

• Experience working with Prometheus, Grafana, and large-scale alerting systems.

• Strong programming skills in Go or Python for the development of operators/CRDs.

• An aptitude for AIOps — viewing the Kubernetes control plane as a platform for automated remediation rather than just a scheduler.

• Experience in integrating an autoscaler/remediator loop into Kubernetes or the capability to design one.

• A runbook-as-code mentality — every SRE playbook you create should be executable by the platform.


🏝️ Benefits

• Competitive salary and performance-based bonuses.

• Comprehensive health, dental, and vision insurance.

• Flexible work hours and remote work options.

• Opportunities for professional development and continuous learning.

• A collaborative and inclusive work environment.

People also viewed

FourEnergy GmbH11 hours ago

Senior DevOps Engineer – Operations

DE flagGermany OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
ICF13 hours ago

Lead DevOps Engineer

US flagVirginia OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$131.3k – $223.1k/year
ApplyView job
Mastercam18 hours ago

DevSecOps Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
C&S Informática22 hours ago

DevOps Engineer – Freelance/Contract, Mid-Level/Senior

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Convene1 day ago

Support and Deployment Engineer

SA flagSaudi Arabia OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Verity Group1 day ago

SRE Engineer

BR flagBrazil OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers