Remotery

Senior DevOps Engineer – Storage

Posted 6 days ago

This is a fully remote position, open to applicants in United States.

📋 Description

• Integrate NFS-based high-performance storage into Kubernetes clusters utilizing CSI, storage classes, and persistent volumes.

• Optimize NFS mount parameters, nconnect/RDMA, Linux client configurations, and network settings for high-throughput, low-latency GPU/AI workloads.

• Deploy and manage storage services and operators; oversee capacity, quotas, snapshots, and lifecycle management.

• Configure and fine-tune Linux systems for storage workloads, focusing on drivers, file systems, networking, and kernel settings.

• Provide storage integration for k0s-based Kubernetes utilizing Cluster API and K0rdent management/child cluster architectures.

• Manage storage operations in fully disconnected air-gapped environments, including Harbor artifact/mirror connectivity and PKI/TLS considerations.

• Automate storage provisioning and configuration using Terraform/OpenTofu and ArgoCD or Flux GitOps pipelines.

• Develop monitoring, alerting, and observability solutions for storage performance, capacity, and health.

• Troubleshoot and resolve performance, reliability, and scaling challenges across the storage stack.

• Establish operational standards and facilitate communication across teams.


⛳️ Requirements

• Over 7 years of experience in Site Reliability Engineering (SRE) or infrastructure operations.

• More than 5 years of experience in building and operating distributed production storage systems at scale.

• Practical experience with high-performance storage solutions such as VAST, Weka, DDN, and PowerScale.

• Fundamental knowledge of Linux and Kubernetes storage, including NFS and CSI.

• Extensive understanding of Linux storage and networking, including kernel and NFS-client layers.

• Experience with infrastructure-as-code and GitOps methodologies.

• Preferred: experience with bare-metal host provisioning, raw disk/hardware layout, and physical server storage configurations.

• Preferred: hands-on experience with VAST and/or Dell PowerScale.

• Preferred: knowledge of GPUDirect Storage and RDMA/RoCE data paths.

• Preferred: familiarity with the Mirantis K0rdent stack, K0rdent Enterprise, K0rdent AI, k0s, MKE, and Cluster API.

• Preferred: experience with Ceph, object/S3 storage backends, and CSI driver operations.

• Preferred: experience in sovereign or high-security air-gapped environments.


🏝️ Benefits

• Opportunities for professional development and training.

• Participation in conferences and working groups.

• Company outings, happy hours, hackathons, and tech talks.

• Competitive compensation package complemented by a robust benefits plan.

• Remote work flexibility.

People also viewed

DATAGROUP2 days ago

DevOps Engineer

DE flagGermany OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Ambush2 days ago

DevOps Engineer

BR flagBrazil OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
DuoKey2 days ago

DevOps Engineer

MU flagMauritius OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
TEKsystems3 days ago

SRE – CloudOps, Practice Architect II

US flagIllinois OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$148.2k – $222.4k/year
ApplyView job
TEKsystems3 days ago

SRE CloudOps Practice Architect II

US flagTexas OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$148.2k – $222.4k/year
ApplyView job
Level Data3 days ago

Senior DevOps Engineer

US flagMassachusetts OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$120k – $135k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers