
Senior DevOps Engineer – Storage
Posted 6 days ago

Posted 6 days ago
This is a fully remote position, open to applicants in United States.
• Integrate NFS-based high-performance storage into Kubernetes clusters utilizing CSI, storage classes, and persistent volumes.
• Optimize NFS mount parameters, nconnect/RDMA, Linux client configurations, and network settings for high-throughput, low-latency GPU/AI workloads.
• Deploy and manage storage services and operators; oversee capacity, quotas, snapshots, and lifecycle management.
• Configure and fine-tune Linux systems for storage workloads, focusing on drivers, file systems, networking, and kernel settings.
• Provide storage integration for k0s-based Kubernetes utilizing Cluster API and K0rdent management/child cluster architectures.
• Manage storage operations in fully disconnected air-gapped environments, including Harbor artifact/mirror connectivity and PKI/TLS considerations.
• Automate storage provisioning and configuration using Terraform/OpenTofu and ArgoCD or Flux GitOps pipelines.
• Develop monitoring, alerting, and observability solutions for storage performance, capacity, and health.
• Troubleshoot and resolve performance, reliability, and scaling challenges across the storage stack.
• Establish operational standards and facilitate communication across teams.
• Over 7 years of experience in Site Reliability Engineering (SRE) or infrastructure operations.
• More than 5 years of experience in building and operating distributed production storage systems at scale.
• Practical experience with high-performance storage solutions such as VAST, Weka, DDN, and PowerScale.
• Fundamental knowledge of Linux and Kubernetes storage, including NFS and CSI.
• Extensive understanding of Linux storage and networking, including kernel and NFS-client layers.
• Experience with infrastructure-as-code and GitOps methodologies.
• Preferred: experience with bare-metal host provisioning, raw disk/hardware layout, and physical server storage configurations.
• Preferred: hands-on experience with VAST and/or Dell PowerScale.
• Preferred: knowledge of GPUDirect Storage and RDMA/RoCE data paths.
• Preferred: familiarity with the Mirantis K0rdent stack, K0rdent Enterprise, K0rdent AI, k0s, MKE, and Cluster API.
• Preferred: experience with Ceph, object/S3 storage backends, and CSI driver operations.
• Preferred: experience in sovereign or high-security air-gapped environments.
• Opportunities for professional development and training.
• Participation in conferences and working groups.
• Company outings, happy hours, hackathons, and tech talks.
• Competitive compensation package complemented by a robust benefits plan.
• Remote work flexibility.
DATAGROUP
Ambush
DuoKey
TEKsystems
Get handpicked remote jobs straight to your inbox weekly.