
Cloud Storage Integration Engineer β Storage, Image, Registry
Posted Aug 28

Posted Aug 28
This is a fully remote position, open to applicants in California, +1 more state.
β’ Design and implement distributed and parallel file systems within the GPU cloud to facilitate AI training and inference I/O patterns.
β’ Take ownership of the complete distributed-storage integration process, which includes provisioning, mounting, multi-tenant isolation, quota management, and lifecycle management within the platform and control plane.
β’ Optimize storage throughput and latency for large-scale parallel access, focusing on dataset loading and checkpointing.
β’ Conduct benchmarks on storage performance across various GPU SKUs and workloads.
β’ Architect multi-region storage solutions that address data locality, replication, consistency, durability, failure domains, and cross-region accessibility.
β’ Manage golden images, templates, GPU drivers/CUDA, and the container/image registry, ensuring versioned releases and multi-region distribution.
β’ Develop monitoring systems, capacity planning strategies, and runbooks to eliminate single points of failure.
β’ Collaborate with the Compute, Network, and Control Plane teams.
β’ Minimum of 3 years of experience in storage engineering or platform infrastructure; for senior-level positions, 6+ years are required.
β’ Practical experience with distributed and parallel file systems.
β’ In-depth knowledge of distributed file system internals, including data and metadata separation, replication, consistency models, and the differences between POSIX and object semantics.
β’ Demonstrated experience in integrating and managing distributed storage in production environments, such as Ceph, Lustre, GPFS/Spectrum Scale, BeeGFS, JuiceFS, or MinIO.
β’ Expertise in performance tuning for high-throughput and parallel I/O operations.
β’ Strong proficiency in Linux systems and automation skills using Python or Go, along with CI/CD practices.
β’ Familiarity with NVMe, RDMA/RoCE storage networking, and caching technologies is highly desirable.
β’ Experience in HPC/AI storage or multi-region storage solutions is a significant advantage.
β’ Competitive salary and comprehensive benefits package.
β’ Opportunities for professional development and career growth.
β’ Collaborative and innovative work environment.
β’ Flexible work arrangements to support work-life balance.
NVIDIA
Cisco
Gateway Ticketing Systems UK Ltd
MoneyGram
Get handpicked remote jobs straight to your inbox weekly.