Senior GPU Cloud Storage Solutions Expert – SRE SME

Posted Sep 11

This is a fully remote position, open to applicants in California, +1 more state.

📋 Description

• Deploy and manage parallel/distributed storage systems such as WEKA, VAST Data, Ceph, and DDN/Lustre.

• Design storage architectures that are tailored to AI workload patterns, including checkpoint I/O bursts, sequential dataset reads, and key-value cache for inference.

• Implement storage isolation for multiple tenants with quality of service (QoS), quotas, and access controls for each tenant.

• Configure and enhance GPU Direct Storage for direct data paths between GPUs and storage.

• Deploy and oversee storage networking solutions encompassing NFS over RDMA, NVMe-oF, high-speed storage fabrics, and Nvidia CMX for comprehensive storage orchestration across clusters.

• Diagnose and optimize storage performance by analyzing IOPS, throughput, latency profiling, fio, IOR, and mdtest.

• Take ownership of the runbook addressing common failure scenarios.

• Plan storage capacity in accordance with GPU cluster expansion and forecasted customer workloads.

• Manage firmware updates, data migration, and disaster recovery protocols.

• Gather storage telemetry data including IO tail latency, checkpoint durations, NVMe SMART data, filesystem health, and RDMA counters.

• Integrate telemetry data into the metrics, logs, and traces repository of the platform team.

• Collaborate with the platform team to establish a storage-fault predictor that includes relevant signals, labels, and acceptable false-positive rates.

• Transform novel incidents into automated solutions, advancing from standard operating procedures (SOPs) to runbook-as-code and agent-executable fixes.

• Provide observability and a foundational predictor for the three primary categories of storage faults.

• Minimize the mean time to recovery (MTTR) for storage incidents.

• Design storage systems for Nvidia GB200-class clusters.


⛳️ Requirements

• A minimum of 5 years of experience in enterprise or high-performance computing (HPC) storage operations, with at least 2 years dedicated to supporting AI/ML workloads.

• Practical experience in deployment and operations with at least two of the following: WEKA, VAST Data, Ceph, DDN/Lustre.

• Strong comprehension of AI training I/O patterns, including checkpoint frequency, dataset loading, and shuffle buffers.

• Familiarity with high-performance storage networking technologies, such as NFS over RDMA and NVMe-oF.

• Understanding of GPU Direct Storage and RDMA-based data transfer methodologies.

• Expertise in storage performance benchmarking and tuning using tools like fio, IOR, and mdtest.

• Experience in implementing multi-tenant storage solutions with isolation and quality of service.

• Solid knowledge of Linux systems, including kernel tuning, filesystem internals, and block device management.

• Experience in delivering an anomaly detector for storage/IO telemetry or the ability to clearly define the required labels and features.

• A runbook-as-code mentality, ensuring every SOP can be executed by a machine within a quarter.


🏝️ Benefits

• Competitive salary and performance-based bonuses.

• Comprehensive health, dental, and vision insurance.

• Opportunities for professional development and training.

• Flexible working hours and remote work options.

• Collaborative and innovative work environment.

People also viewed

FourEnergy GmbH11 hours ago

Senior DevOps Engineer – Operations

DE flagGermany OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
ICF13 hours ago

Lead DevOps Engineer

US flagVirginia OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$131.3k – $223.1k/year
ApplyView job
Mastercam17 hours ago

DevSecOps Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
C&S Informática22 hours ago

DevOps Engineer – Freelance/Contract, Mid-Level/Senior

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Convene1 day ago

Support and Deployment Engineer

SA flagSaudi Arabia OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Verity Group1 day ago

SRE Engineer

BR flagBrazil OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers