Remotery

Senior DevOps Engineer – Storage

Posted 2 days ago

This is a fully remote position, open to applicants in United States.

📋 Description

• Deploy, integrate, and manage high-performance storage solutions for GPU-accelerated computing and AI platforms.

• Implement NFS-based high-performance storage within Kubernetes clusters through CSI, storage classes, and persistent volumes.

• Optimize NFS mount options, nconnect/RDMA, Linux client configurations, and network settings to support high-throughput, low-latency GPU/AI workloads.

• Deploy and manage storage services and operators; oversee capacity, quotas, snapshots, and lifecycle management.

• Configure and enhance Linux systems for storage workloads, focusing on drivers, file system layouts, network tuning, and kernel parameters.

• Provide storage integration for k0s-based Kubernetes utilizing Cluster API (CAPI) and K0rdent management/child cluster architectures.

• Manage storage in completely disconnected air-gapped settings, including Harbor artifact/mirror connectivity and PKI/TLS considerations.

• Automate storage provisioning and setup using Terraform/OpenTofu and ArgoCD or Flux GitOps pipelines.

• Develop monitoring, alerting, and observability systems for storage performance, capacity, and health metrics.

• Troubleshoot and resolve performance, reliability, and scalability challenges across the storage ecosystem.

• Establish operational standards and facilitate communication across teams.


⛳️ Requirements

• Over 7 years of experience in Site Reliability Engineering (SRE) or infrastructure operations.

• More than 5 years of experience in building and operating distributed production storage systems at scale.

• Practical experience with high-performance storage technologies such as VAST, Weka, DDN, and PowerScale.

• Understanding of Linux and Kubernetes storage fundamentals.

• Familiar with NFS and Container Storage Interface (CSI).

• Strong knowledge of Linux storage and networking fundamentals, including kernel and NFS-client layers.

• Experience with infrastructure-as-code practices and GitOps methodologies.

• Proficient with Terraform/OpenTofu and ArgoCD or Flux.

• Background in Kubernetes, k0s, Cluster API (CAPI), and K0rdent.

• Experience working in hybrid, edge, and fully disconnected air-gapped environments.

• Knowledge of Harbor and PKI/TLS considerations.

• Experience with bare-metal hardware is a significant advantage.


🏝️ Benefits

• Opportunities for professional development and training.

• Participation in conferences and working groups.

• Company events, happy hours, hackathons, and tech talks.

• Competitive salary package with an extensive benefits plan.

• Flexibility for remote work.

People also viewed

Ontrac Solutions2 days ago

Site Reliability Engineer

PK flagPakistan OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
CyberSheath2 days ago

Cloud Operations Engineer

US flagVirginia OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$110k – $127k/year
ApplyView job
Ontrac Solutions2 days ago

Site Reliability Engineer

IN flagIndia OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
NVIDIA2 days ago

Service Reliability Engineer

US flagTexas OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$168k – $333.5k/year
ApplyView job
Nagarro2 days ago

Senior Site Reliability Engineer, AWS Cloud

RO flagRomania OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Capgemini2 days ago

Senior DevOps Engineer

UA flagUkraine OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers