
AI Storage Solutions Expert
Posted 21 hours ago

Posted 21 hours ago
This is a fully remote position, open to applicants in California, +1 more state.
• Deploy and manage parallel/distributed storage systems such as WEKA, VAST Data, Ceph, and DDN/Lustre.
• Craft storage architectures tailored for AI checkpoint I/O surges, sequential dataset retrievals, and inference KV caching.
• Enforce multi-tenant storage isolation featuring per-tenant quality of service, quotas, and access management.
• Configure and enhance GPU Direct Storage for seamless GPU-to-storage data pathways.
• Set up and oversee storage networking, encompassing NFS over RDMA, NVMe-oF, high-speed storage fabrics, and Nvidia CMX.
• Analyze and optimize storage performance via IOPS, throughput, latency profiling, fio, IOR, and mdtest.
• Maintain runbooks for typical storage failure scenarios.
• Strategize storage capacity planning for GPU cluster expansion and customer workload forecasts.
• Oversee firmware updates, data migration, and disaster recovery protocols.
• Implement storage telemetry and integrate IO tail latency, checkpoint durations, NVMe SMART metrics, filesystem health, and RDMA counters into the platform’s metrics/logs/traces repository.
• Collaborate with the platform team to establish the storage-fault predictor, including identifying signals, incident-derived labels, and acceptable false-positive rates.
• Transform new incidents into automated solutions, evolving from standard operating procedures to runbook-as-code and agent-executable remediation.
• Provide observability and establish a baseline predictor for the top three storage-fault categories.
• Design storage systems for Nvidia GB200-class cluster implementations.
• Minimize mean time to recovery (MTTR) for storage incidents.
• Minimum of 5 years of experience in enterprise or HPC storage operations.
• At least 2 years of experience in supporting AI/ML workloads.
• Practical deployment and operational experience with a minimum of two from WEKA, VAST Data, Ceph, and DDN/Lustre.
• In-depth understanding of AI training I/O behaviors, including checkpoint frequencies, dataset loading, and shuffle buffers.
• Experience with high-performance storage networking solutions, including NFS over RDMA and NVMe-oF.
• Familiarity with GPU Direct Storage and RDMA-driven data transfers.
• Expertise in storage performance benchmarking and tuning, utilizing tools like fio, IOR, and mdtest.
• Experience in implementing multi-tenant storage with isolation and quality of service features.
• Strong Linux systems expertise, covering kernel tuning, filesystem internals, and block device management.
• Experience in deploying an anomaly detector for storage/IO telemetry, or the ability to define necessary labels and features.
• A runbook-as-code mentality; standard operating procedures should be executable by a machine within a quarter.
• Competitive salary and performance-based incentives.
• Comprehensive health, dental, and vision insurance plans.
• Opportunities for professional development and continuous learning.
• Flexible working hours and remote work options.
• Engaging company culture with team-building activities and events.
Progressive Leasing
apna
apna
Texas Research International
Get handpicked remote jobs straight to your inbox weekly.