Remotery

K8 Site Reliability SME

Posted 20 hours ago

This is a fully remote position, open to applicants in California, +1 more state.

📋 Description

• Design, implement, and manage production Kubernetes control planes that are optimized for large-scale GPU workloads (ranging from 100 to 10,000 GPUs).

• Configure the Nvidia GPU Operator, device plugins, MIG, and policies for GPU time-slicing.

• Execute topology-aware scheduling that considers GPU locality, NVLink domain awareness, and network rail affinity.

• Create Custom Resource Definitions (CRDs) to manage the lifecycle of GPU workloads.

• Integrate Slurm, Ray, and Kubeflow within Kubernetes environments.

• Ensure multi-tenant isolation through the use of namespaces, network policies, resource quotas, RBAC, and pod security standards.

• Automate the provisioning of Bare-Metal-as-a-Service, including tenant onboarding, lifecycle management, and reclamation.

• Develop Terraform providers and modules to facilitate infrastructure-as-code across GPU clusters.

• Define and publish Service Level Indicators (SLIs) and Service Level Objectives (SLOs) for cluster availability, job completion rates, and provisioning latency.

• Automate incident management runbooks, escalation processes, and post-incident reviews.

• Operate monitoring stacks including Prometheus, Grafana, Alertmanager, and PagerDuty.

• Automate the detection of GPU node failures, including drain, cordon, taint, and workload rescheduling.

• Ensure the Kubernetes control plane is secure for automated AIOps remediation.

• Provide CRD schemas for platform predictors and remediation tools.

• Transform human interventions into automated workflows.

• Enable automated draining and rescheduling in response to predicted GPU faults without impacting customers.

• Launch Bare-Metal-as-a-Service for external tenants with self-service onboarding capabilities.

• Achieve cluster availability and job completion SLOs.


⛳️ Requirements

• A minimum of 5 years in Kubernetes operations, with at least 2 years focused on managing GPU workloads in Kubernetes.

• In-depth knowledge of the Nvidia GPU Operator, device plugins, and GPU scheduling within Kubernetes.

• Experience with topology-aware scheduling and the management of GPU-specific resources.

• Practical experience in building multi-tenant Kubernetes platforms with robust isolation guarantees.

• Familiarity with bare-metal server provisioning and lifecycle automation using Ironic, MAAS, or custom solutions.

• Expertise in Terraform, Helm, and GitOps workflows such as ArgoCD or Flux.

• Strong background in Site Reliability Engineering (SRE), including frameworks for SLIs/SLOs, incident management, and capacity planning.

• Experience with Prometheus, Grafana, and large-scale alerting systems.

• Proficient programming skills in Go or Python for developing operators and CRDs.

• Aptitude for AIOps: the capability to treat the Kubernetes control plane as a platform for automated remediation.

• Experience in integrating an autoscaler/remediator loop into Kubernetes, or the ability to design such a system.

• A mindset geared towards runbook-as-code practices.

• Compliance with equal employment opportunity regulations as per relevant country, state, and local laws.


🏝️ Benefits

• Comprehensive health insurance coverage.

• Flexible work hours and remote work options.

• Opportunities for professional development and continuous learning.

• Collaborative and innovative work environment.

People also viewed

CVS Health6 hours ago

Salesforce DevOps Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$83.4k – $166.9k/year
ApplyView job
Devoteam7 hours ago

Data, AWS DevSecOps

ES flagSpain OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Aspirion7 hours ago

Senior DevOps Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Goodgame Studios7 hours ago

Senior Agentic Engineer – Java Backend, DevOps

DE flagGermany OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Instacart8 hours ago

Site Reliability Engineer II

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$133k – $169k/year
ApplyView job
Logicalis Spain8 hours ago

DevOps Engineer

ES flagSpain OnlyFull-timeDevOps & Site Reliability Engineer (SRE)€40k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers