
Senior SRE, Automation Engineer β Customer Facing
Posted Aug 28

Posted Aug 28
This is a fully remote position, open to applicants in California, +1 more state.
β’ Take full responsibility for the reliability of the customer-facing GPU cloud service, including aspects such as availability, job completion, provisioning latency, and tenant experience.
β’ Manage production Kubernetes clusters optimized for GPU workloads, ranging from 100 to 10,000 GPUs.
β’ Oversee the Nvidia GPU operator, device plugin, MIG configuration, GPU time-slicing, and policies for multi-tenant GPU allocation.
β’ Implement scheduling that is aware of topology, considering GPU locality, NVLink domains, and network rail affinity.
β’ Handle tenant onboarding, quota management, enforcement of isolation, and offboarding/reclamation processes.
β’ Automate the provisioning of Bare-Metal-as-a-Service, tenant handoffs, lifecycle management, and reclamation tasks.
β’ Define and publish SLIs, SLOs, and SLAs that are customer-facing.
β’ Lead the detection of incidents, their remediation, escalation processes, customer-facing status updates, and post-incident reviews.
β’ Develop and maintain tenant-aware monitoring and observability tools using Prometheus, Grafana, Alertmanager, and PagerDuty.
β’ Automate the detection of GPU node failures, as well as actions such as drain/cordon/taint and rescheduling of workloads.
β’ Develop infrastructure as code, leveraging tools such as Terraform, Helm, and GitOps across GPU clusters.
β’ Create and maintain service documentation, tenant runbooks, capacity plans, and support tiering structures.
β’ Collaborate with customer success and support teams to translate customer-reported issues into systemic enhancements.
β’ Construct self-service observability features for customers regarding their status, quotas, and job health.
β’ Ensure the control plane is secure for automated remediation through the use of CRDs and executable runbooks.
β’ Achieve published and met SLAs, automate GPU fault recovery, facilitate external-tenant BMaaS onboarding, and reduce MTTD/MTTR, while providing tenant self-service observability.
β’ Minimum of 5 years' experience in SRE or cloud operations, with at least 2 years focused on managing GPU workloads at scale.
β’ Profound understanding of Kubernetes operations and GPU workload management, including Nvidia GPU operator, device plugin, MIG, time-slicing, and GPU scheduling.
β’ Experience in topology-aware scheduling and management of GPU-specific resources.
β’ Practical experience in building multi-tenant cloud platforms that offer robust isolation guarantees.
β’ Customer-facing cloud service experience, with skills in defining and managing SLAs/SLOs and handling tenant incidents and communications.
β’ Familiarity with bare-metal server provisioning and lifecycle automation using Ironic, MAAS, or custom solutions.
β’ Proficient in Terraform, Helm, and GitOps workflows, with experience in ArgoCD or Flux.
β’ Strong background in SRE, including familiarity with SLI/SLO/SLA frameworks, error budgets, incident management, and capacity planning.
β’ Experience with Prometheus, Grafana, and large-scale alerting systems.
β’ Strong programming skills in Go or Python for the development of automation and operators.
β’ Aptitude for AIOps, treating the control plane as an execution surface for automated remediation.
β’ Embrace a runbook-as-code mindset, where SRE playbooks are designed to be executable by the platform.
β’ Competitive salary and comprehensive benefits package.
β’ Opportunity for professional growth and development.
β’ Flexible work environment with remote work options.
β’ Engaging and innovative team culture.
Koniag Government Services
LMI
AmeriSave Mortgage Corporation
Moneycorp
Get handpicked remote jobs straight to your inbox weekly.