
K8 Site Reliability SME
Posted 20 hours ago

Posted 20 hours ago
This is a fully remote position, open to applicants in California, +1 more state.
• Design, implement, and manage production Kubernetes control planes that are optimized for large-scale GPU workloads (ranging from 100 to 10,000 GPUs).
• Configure the Nvidia GPU Operator, device plugins, MIG, and policies for GPU time-slicing.
• Execute topology-aware scheduling that considers GPU locality, NVLink domain awareness, and network rail affinity.
• Create Custom Resource Definitions (CRDs) to manage the lifecycle of GPU workloads.
• Integrate Slurm, Ray, and Kubeflow within Kubernetes environments.
• Ensure multi-tenant isolation through the use of namespaces, network policies, resource quotas, RBAC, and pod security standards.
• Automate the provisioning of Bare-Metal-as-a-Service, including tenant onboarding, lifecycle management, and reclamation.
• Develop Terraform providers and modules to facilitate infrastructure-as-code across GPU clusters.
• Define and publish Service Level Indicators (SLIs) and Service Level Objectives (SLOs) for cluster availability, job completion rates, and provisioning latency.
• Automate incident management runbooks, escalation processes, and post-incident reviews.
• Operate monitoring stacks including Prometheus, Grafana, Alertmanager, and PagerDuty.
• Automate the detection of GPU node failures, including drain, cordon, taint, and workload rescheduling.
• Ensure the Kubernetes control plane is secure for automated AIOps remediation.
• Provide CRD schemas for platform predictors and remediation tools.
• Transform human interventions into automated workflows.
• Enable automated draining and rescheduling in response to predicted GPU faults without impacting customers.
• Launch Bare-Metal-as-a-Service for external tenants with self-service onboarding capabilities.
• Achieve cluster availability and job completion SLOs.
• A minimum of 5 years in Kubernetes operations, with at least 2 years focused on managing GPU workloads in Kubernetes.
• In-depth knowledge of the Nvidia GPU Operator, device plugins, and GPU scheduling within Kubernetes.
• Experience with topology-aware scheduling and the management of GPU-specific resources.
• Practical experience in building multi-tenant Kubernetes platforms with robust isolation guarantees.
• Familiarity with bare-metal server provisioning and lifecycle automation using Ironic, MAAS, or custom solutions.
• Expertise in Terraform, Helm, and GitOps workflows such as ArgoCD or Flux.
• Strong background in Site Reliability Engineering (SRE), including frameworks for SLIs/SLOs, incident management, and capacity planning.
• Experience with Prometheus, Grafana, and large-scale alerting systems.
• Proficient programming skills in Go or Python for developing operators and CRDs.
• Aptitude for AIOps: the capability to treat the Kubernetes control plane as a platform for automated remediation.
• Experience in integrating an autoscaler/remediator loop into Kubernetes, or the ability to design such a system.
• A mindset geared towards runbook-as-code practices.
• Compliance with equal employment opportunity regulations as per relevant country, state, and local laws.
• Comprehensive health insurance coverage.
• Flexible work hours and remote work options.
• Opportunities for professional development and continuous learning.
• Collaborative and innovative work environment.
CVS Health
Devoteam
Aspirion
Goodgame Studios
Get handpicked remote jobs straight to your inbox weekly.