
Senior Software Systems Engineer – Storage
Posted Aug 27

Posted Aug 27
This is a fully remote position, open to applicants in United States.
• Integrate high-performance NFS-based storage solutions, such as VAST and Dell PowerScale, into Kubernetes clusters utilizing CSI, storage classes, and persistent volumes.
• Optimize NFS mount options, nconnect/RDMA, Linux clients, and network configurations to support high-throughput, low-latency GPU/AI workloads.
• Deploy and manage storage services and operators efficiently.
• Oversee storage capacity, quotas, snapshots, and lifecycle management.
• Configure and enhance Linux systems for storage workloads, focusing on drivers, file-system structure, network optimization, and kernel parameters.
• Provide storage integration for k0s-based Kubernetes using Cluster API and K0rdent management/child-cluster architectures.
• Manage storage operations in completely disconnected air-gapped environments, addressing Harbor artifact/mirror connectivity and PKI/TLS requirements.
• Automate storage provisioning and configuration with Terraform/OpenTofu and GitOps pipelines using ArgoCD or Flux.
• Develop monitoring, alerting, and observability solutions for storage performance, capacity, and health.
• Troubleshoot and resolve performance, reliability, and scaling challenges throughout the storage stack.
• Establish operational standards and facilitate communication across teams.
• A minimum of 7 years of experience in Site Reliability Engineering (SRE) or infrastructure operations.
• At least 5 years of experience in building and managing distributed production storage systems at scale.
• Practical experience with high-performance storage solutions such as VAST, Weka, DDN, or PowerScale.
• Solid understanding of Linux and Kubernetes storage fundamentals, including NFS and CSI.
• Extensive knowledge of Linux storage and networking, specifically in kernel and NFS-client layers.
• Experience with infrastructure-as-code and GitOps methodologies.
• Proven ability to diagnose performance and reliability issues from end to end.
• Preference for candidates with bare-metal hardware experience.
• Hands-on experience with VAST and/or Dell PowerScale is highly desirable.
• Familiarity with GPUDirect Storage and RDMA/RoCE data paths is preferred.
• Experience with the Mirantis K0rdent stack, k0s, MKE, and Cluster API is advantageous.
• Knowledge of Ceph, object/S3 storage, and CSI driver operations is preferred.
• Experience in sovereign or high-security air-gapped environments is a plus.
• Opportunities for professional development and training.
• Participation in conferences and collaborative working groups.
• Company outings, happy hours, hackathons, and tech talks.
• Competitive compensation package complemented by a robust benefits plan.
• Flexible remote work arrangement.
Kodama Systems
Latitude IT Solutions | SDVOSB
MyFitnessPal
GE Vernova
Get handpicked remote jobs straight to your inbox weekly.