
Senior Storage Software Engineer β DGX Cloud
Posted 6 days ago

Posted 6 days ago
This is a fully remote position, open to applicants in California.
β’ Contribute code to open-source parallel and distributed file systems, as well as distributed object storage.
β’ Engage with upstream communities and maintainers by upstreaming fixes and features.
β’ Serve as a hands-on storage software lead, responsible for writing and reviewing production code.
β’ Diagnose bugs by analyzing kernel, NFS, NVMe-oF, or SPDK source code.
β’ Make critical technical decisions on storage deliveries based on measurable targets.
β’ Triage, troubleshoot, and root-cause complex storage issues across extensive GPU clusters.
β’ Investigate I/O and metadata performance, as well as issues related to data corruption and recovery.
β’ Validate the architecture, capabilities, performance, and durability of storage solutions.
β’ Conduct scale tests, benchmarks, and recovery drills to ensure system reliability.
β’ Qualify new builds according to defined performance and durability metrics.
β’ Define and recommend best practices for configuration, tuning, and operations of high-performance file systems on GPU infrastructure.
β’ Assist operators and internal customers in applying storage guidelines effectively.
β’ Collaborate with teams in training, inference, accelerated computing, SRE, operations, networking, and security.
β’ Work alongside cloud providers, neocloud operators, and storage vendors to develop common architecture.
β’ Utilize modern AI coding and agentic tools to enhance building, debugging, validation, and operational processes.
β’ BS, MS, or PhD in Computer Science, Electrical Engineering, or a related field β or equivalent experience.
β’ More than 12 years of direct experience in storage software engineering.
β’ Extensive experience with high-performance parallel or distributed file systems managing multi-petabyte scale.
β’ Contributions to open-source projects that involve distributed or parallel file systems.
β’ Hands-on experience in writing and reviewing production code, examining file system, kernel, NVMe-oF, or SPDK source, and conducting scale tests or recovery drills.
β’ Proven ability to diagnose and resolve storage challenges in large GPU or HPC clusters, including I/O and metadata performance analysis.
β’ Strong expertise in at least one systems programming language: C, C++, Rust, or Go.
β’ Proficiency in Python programming.
β’ Comfortable working with Linux kernel storage and networking stacks, including block layer, RDMA/RoCE/InfiniBand, NVMe, page cache, VFS, and multipath configurations.
β’ Solid understanding of object storage concepts, including S3/Swift-class implementations.
β’ Strong grasp of block storage technologies, including NVMe-oF and iSCSI.
β’ Excellent written and verbal communication skills.
β’ Ability to operate in a 24/7 production environment.
β’ Security-first mindset.
β’ Maintainers or contributors to widely used public projects.
β’ Experience in crafting or operating storage solutions for AI training or inference at a large GPU scale.
β’ Background in kernel and file system development, metadata scalability, data placement, failure recovery, or equivalent experience.
β’ Experience with Kubernetes and CSI driver development for storage solutions.
β’ Hands-on experience with SPDK, libfabric, or FUSE performance optimization techniques.
β’ Equity.
β’ Benefits.
Shield AI
Netflix
Travoom
HeroSoftware GmbH - Shopify Apps
Get handpicked remote jobs straight to your inbox weekly.