
Senior GPU Cloud Storage Solutions Expert – SRE SME
Posted Sep 11

Posted Sep 11
This is a fully remote position, open to applicants in California, +1 more state.
• Deploy and manage parallel/distributed storage systems such as WEKA, VAST Data, Ceph, and DDN/Lustre.
• Design storage architectures that are tailored to AI workload patterns, including checkpoint I/O bursts, sequential dataset reads, and key-value cache for inference.
• Implement storage isolation for multiple tenants with quality of service (QoS), quotas, and access controls for each tenant.
• Configure and enhance GPU Direct Storage for direct data paths between GPUs and storage.
• Deploy and oversee storage networking solutions encompassing NFS over RDMA, NVMe-oF, high-speed storage fabrics, and Nvidia CMX for comprehensive storage orchestration across clusters.
• Diagnose and optimize storage performance by analyzing IOPS, throughput, latency profiling, fio, IOR, and mdtest.
• Take ownership of the runbook addressing common failure scenarios.
• Plan storage capacity in accordance with GPU cluster expansion and forecasted customer workloads.
• Manage firmware updates, data migration, and disaster recovery protocols.
• Gather storage telemetry data including IO tail latency, checkpoint durations, NVMe SMART data, filesystem health, and RDMA counters.
• Integrate telemetry data into the metrics, logs, and traces repository of the platform team.
• Collaborate with the platform team to establish a storage-fault predictor that includes relevant signals, labels, and acceptable false-positive rates.
• Transform novel incidents into automated solutions, advancing from standard operating procedures (SOPs) to runbook-as-code and agent-executable fixes.
• Provide observability and a foundational predictor for the three primary categories of storage faults.
• Minimize the mean time to recovery (MTTR) for storage incidents.
• Design storage systems for Nvidia GB200-class clusters.
• A minimum of 5 years of experience in enterprise or high-performance computing (HPC) storage operations, with at least 2 years dedicated to supporting AI/ML workloads.
• Practical experience in deployment and operations with at least two of the following: WEKA, VAST Data, Ceph, DDN/Lustre.
• Strong comprehension of AI training I/O patterns, including checkpoint frequency, dataset loading, and shuffle buffers.
• Familiarity with high-performance storage networking technologies, such as NFS over RDMA and NVMe-oF.
• Understanding of GPU Direct Storage and RDMA-based data transfer methodologies.
• Expertise in storage performance benchmarking and tuning using tools like fio, IOR, and mdtest.
• Experience in implementing multi-tenant storage solutions with isolation and quality of service.
• Solid knowledge of Linux systems, including kernel tuning, filesystem internals, and block device management.
• Experience in delivering an anomaly detector for storage/IO telemetry or the ability to clearly define the required labels and features.
• A runbook-as-code mentality, ensuring every SOP can be executed by a machine within a quarter.
• Competitive salary and performance-based bonuses.
• Comprehensive health, dental, and vision insurance.
• Opportunities for professional development and training.
• Flexible working hours and remote work options.
• Collaborative and innovative work environment.
FourEnergy GmbH
ICF
Mastercam
C&S Informática
Get handpicked remote jobs straight to your inbox weekly.