
Staff Storage Platform Engineer β AI Storage
Posted Sep 8

Posted Sep 8
This is a fully remote position, open to applicants in United States.
β’ Design, develop, and manage the AI storage layer that supports extensive GPU infrastructure.
β’ Architect scalable storage solutions for both edge and core deployments.
β’ Establish storage strategies for distributed inference, fine-tuning, and training workloads.
β’ Create solutions utilizing hyperconverged storage technologies like StorPool, local NVMe, and disaggregated systems such as VAST and Weka.
β’ Define reference architectures, design principles, reusable patterns, and long-term storage strategies.
β’ Enhance throughput, latency, data locality, resilience, cost efficiency, and operability for GPU-intensive clusters.
β’ Benchmark and verify storage performance under realistic AI workloads.
β’ Implement and sustain CSI drivers, integrating storage with Kubernetes and orchestration systems.
β’ Incorporate block, object, and shared file storage into the platform.
β’ Develop multi-tenant, multi-cluster, and multi-site storage environments.
β’ Create storage architectures for distributed inference platforms such as NVIDIA Dynamo and llm-d.
β’ Optimize KV-cache persistence, token generation pipelines, and high-concurrency inference workloads.
β’ Design high-performance data paths using GPU Direct Storage, RDMA/RoCE, and NVMe-oF.
β’ Lead investigations into storage performance, incident responses, and root-cause analyses.
β’ Enhance reliability, durability, observability, recovery processes, and operational standards.
β’ Oversee the comprehensive delivery of storage infrastructure, from architecture and validation to production rollout.
β’ Drive capacity planning, scaling strategies, lifecycle decisions, and safe production modifications.
β’ Collaborate with compute, networking, platform, DevOps, operations, deployment teams, vendors, and stakeholders.
β’ Serve as the primary authority on storage design and mentor engineers.
β’ Extensive hands-on experience in designing and operating distributed storage systems for high-performance computing environments.
β’ Demonstrated experience in crafting storage architectures for large-scale AI inference or training platforms, covering dataset distribution, checkpointing, and KV-cache storage methodologies.
β’ In-depth understanding of the Linux storage and I/O stack.
β’ Strong grasp of data access patterns in AI workloads.
β’ Experience in optimizing storage for GPU-accelerated tasks.
β’ Familiarity with WEKA Data Platform.
β’ Knowledge of Kubernetes storage integrations, particularly CSI.
β’ Experience in managing large-scale storage clusters.
β’ Profound expertise in designing and operating storage platforms tailored for GPU-heavy environments and distributed AI workloads.
β’ Acquainted with storage hardware, NVMe devices, storage fabrics, and high-performance data pathways.
β’ Experience in designing storage observability systems.
β’ Strong automation capabilities using Python and/or Bash.
β’ Experience implementing software engineering practices in storage automation and operational tools.
β’ Proven track record of leading complex technical projects across multiple teams.
β’ Familiarity with Kubernetes, distributed file systems, object storage, block storage, StorPool, VAST Data, Weka, GPU Direct Storage, RDMA/RoCE, NVMe-oF, and SPDK.
β’ Competitive compensation package that reflects your expertise and experience.
β’ A welcoming work environment defined by friendliness, international diversity, flexibility, and a hybrid-friendly approach.
β’ Opportunity for exciting career advancement in a rapidly growing scale-up.
β’ Equal opportunity employer dedicated to diversity and inclusion.
OPENDataJobs
Presidio
EasyLlama - HR & Compliance Training For Modern Teams
Coinbase
Get handpicked remote jobs straight to your inbox weekly.