
HPC Storage Engineer
Posted Sep 10

Posted Sep 10
This is a fully remote position, open to applicants in United States.
• Take responsibility for the capacity, durability, availability, and performance of network volumes, local NVMe, and S3-compatible object storage.
• Optimize device and filesystem settings, including caching, read-ahead, replication, erasure coding, and client-side mount configurations.
• Identify and resolve intricate performance issues from start to finish.
• Manage capacity expansions, hardware upgrades, migrations, and rebalances without any disruption visible to customers.
• Design and optimize high-throughput storage network paths, taking into consideration MTU, jumbo frames, congestion and flow control, multipath, and NIC/offload configurations.
• Enhance RDMA/RoCE and high-speed IB/Ethernet fabrics for storage traffic.
• Work in collaboration with network engineering on topology, oversubscription, and cross-region data transportation.
• Develop production code using Go, Python, or similar languages for control-plane services, provisioning, data transfer, and monitoring.
• Construct and extend control-plane, S3-compatible, CSI, Kubernetes, vendor, and cloud-provider APIs.
• Automate manual storage processes and manage infrastructure through code.
• Engage in code reviews, testing, and continuous integration activities.
• Implement monitoring for IOPS, throughput, latency, errors, retries, utilization, and consumption per tenant.
• Create dashboards, Service Level Objectives (SLOs), and alerts.
• Take part in storage on-call rotations and lead post-incident reviews with a focus on no-blame culture.
• Assist in determining distributed storage systems, data tiering and placement strategies, network optimization, and large-scale purchasing and deployment of petabyte storage solutions.
• Over 8 years of experience in infrastructure, storage, or systems engineering, with significant responsibility for production storage at scale.
• Extensive, hands-on experience with at least one distributed storage system, such as Ceph, MinIO, Lustre, GPFS/Spectrum Scale, MooseFS, WekaFS, VAST, ZFS-based systems, or similar.
• Strong understanding of Linux internals and the storage stack, including block layer, filesystems, NVMe, page cache, I/O schedulers, NFS/SMB, iSCSI/NVMe-oF.
• Proven experience in building and/or managing S3-compatible object storage services.
• Solid grounding in networking principles, particularly in tuning networks for storage workloads.
• Proficient in writing and deploying production code in Go, Python, Rust, or related languages.
• Hands-on experience with observability tools like Prometheus, Grafana, or Datadog, including metric design.
• A history of performance analysis and debugging under real production stress.
• Self-motivated with a clear direction.
• Committed to continuous improvement.
• Demonstrated ownership across team boundaries.
• Collaborative and humble, yet confident.
• Must be eligible to work in the United States.
• Not requiring employment visa sponsorship.
• Significant equity; all team members receive stock options.
• Comprehensive medical, dental, and vision plans.
• Flexible paid time off (PTO).
• Remote-first work culture.
• Internal communication primarily through Slack.
• A passionate team at the forefront of AI infrastructure.
• $1,200 stipend for home office and equipment.
RR Donnelley
plotdesk
CmdScale GmbH
Colsubsidio
Get handpicked remote jobs straight to your inbox weekly.