Staff Storage Platform Engineer – AI Storage

Posted Sep 8

This is a fully remote position, open to applicants in United States.

πŸ“‹ Description

β€’ Design, develop, and manage the AI storage layer that supports extensive GPU infrastructure.

β€’ Architect scalable storage solutions for both edge and core deployments.

β€’ Establish storage strategies for distributed inference, fine-tuning, and training workloads.

β€’ Create solutions utilizing hyperconverged storage technologies like StorPool, local NVMe, and disaggregated systems such as VAST and Weka.

β€’ Define reference architectures, design principles, reusable patterns, and long-term storage strategies.

β€’ Enhance throughput, latency, data locality, resilience, cost efficiency, and operability for GPU-intensive clusters.

β€’ Benchmark and verify storage performance under realistic AI workloads.

β€’ Implement and sustain CSI drivers, integrating storage with Kubernetes and orchestration systems.

β€’ Incorporate block, object, and shared file storage into the platform.

β€’ Develop multi-tenant, multi-cluster, and multi-site storage environments.

β€’ Create storage architectures for distributed inference platforms such as NVIDIA Dynamo and llm-d.

β€’ Optimize KV-cache persistence, token generation pipelines, and high-concurrency inference workloads.

β€’ Design high-performance data paths using GPU Direct Storage, RDMA/RoCE, and NVMe-oF.

β€’ Lead investigations into storage performance, incident responses, and root-cause analyses.

β€’ Enhance reliability, durability, observability, recovery processes, and operational standards.

β€’ Oversee the comprehensive delivery of storage infrastructure, from architecture and validation to production rollout.

β€’ Drive capacity planning, scaling strategies, lifecycle decisions, and safe production modifications.

β€’ Collaborate with compute, networking, platform, DevOps, operations, deployment teams, vendors, and stakeholders.

β€’ Serve as the primary authority on storage design and mentor engineers.


⛳️ Requirements

β€’ Extensive hands-on experience in designing and operating distributed storage systems for high-performance computing environments.

β€’ Demonstrated experience in crafting storage architectures for large-scale AI inference or training platforms, covering dataset distribution, checkpointing, and KV-cache storage methodologies.

β€’ In-depth understanding of the Linux storage and I/O stack.

β€’ Strong grasp of data access patterns in AI workloads.

β€’ Experience in optimizing storage for GPU-accelerated tasks.

β€’ Familiarity with WEKA Data Platform.

β€’ Knowledge of Kubernetes storage integrations, particularly CSI.

β€’ Experience in managing large-scale storage clusters.

β€’ Profound expertise in designing and operating storage platforms tailored for GPU-heavy environments and distributed AI workloads.

β€’ Acquainted with storage hardware, NVMe devices, storage fabrics, and high-performance data pathways.

β€’ Experience in designing storage observability systems.

β€’ Strong automation capabilities using Python and/or Bash.

β€’ Experience implementing software engineering practices in storage automation and operational tools.

β€’ Proven track record of leading complex technical projects across multiple teams.

β€’ Familiarity with Kubernetes, distributed file systems, object storage, block storage, StorPool, VAST Data, Weka, GPU Direct Storage, RDMA/RoCE, NVMe-oF, and SPDK.


🏝️ Benefits

β€’ Competitive compensation package that reflects your expertise and experience.

β€’ A welcoming work environment defined by friendliness, international diversity, flexibility, and a hybrid-friendly approach.

β€’ Opportunity for exciting career advancement in a rapidly growing scale-up.

β€’ Equal opportunity employer dedicated to diversity and inclusion.

People also viewed

OPENDataJobs1 day ago

Databricks Platform Engineer – DevSecOps

US flagUnited States OnlyFull-timePlatform Engineer$120k – $140k/year
ApplyView job
Presidio1 day ago

Engineer, Platform Tooling

US flagUnited States OnlyFull-timePlatform Engineer
ApplyView job
EasyLlama - HR & Compliance Training For Modern Teams1 day ago

Founding Platform Engineer – Ruby on Rails

US flagCalifornia, +19 more statesFull-timePlatform Engineer$163k – $185k/year
ApplyView job
Coinbase1 day ago

Staff Software Engineer – Platform, Financial Engineering

US flagUnited States OnlyFull-timePlatform Engineer$218k – $256.5k/year
ApplyView job
Stitch Fix1 day ago

Platform Engineer

US flagUnited States OnlyFull-timePlatform Engineer$92.6k – $154.5k/year
ApplyView job
Ambry Genetics1 day ago

Senior Software Engineer – Platform Engineering

US flagCalifornia OnlyFull-timePlatform Engineer$130k – $150k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers