
AI Storage Infrastructure Engineer
Posted Sep 11

Posted Sep 11
This is a fully remote position, open to applicants in United States.
• Integrate high-performance NFS-based storage solutions such as VAST and Dell PowerScale into Kubernetes clusters utilizing CSI, storage classes, and persistent volumes.
• Optimize NFS data paths, including mount options, nconnect/RDMA, Linux clients, and network settings, to support high-throughput, low-latency GPU/AI workloads.
• Deploy and manage storage services and operators effectively.
• Oversee storage capacity, quotas, snapshots, and their lifecycle management.
• Configure and enhance Linux systems for storage workloads, focusing on drivers, file-system layout, network tuning, and kernel parameters.
• Provide storage integration for k0s-based Kubernetes through Cluster API and K0rdent management/child-cluster architectures.
• Operate storage solutions in fully disconnected air-gapped environments, including managing Harbor artifact/mirror connectivity and PKI/TLS aspects.
• Automate storage provisioning and configuration using Terraform/OpenTofu and GitOps pipelines with ArgoCD or Flux.
• Develop monitoring, alerting, and observability systems for storage performance, capacity, and health.
• Troubleshoot and resolve performance, reliability, and scaling challenges throughout the storage stack.
• Establish operational standards and facilitate communication across teams.
• Over 7 years of experience in Site Reliability Engineering (SRE) or hardware/storage infrastructure operations.
• More than 5 years of experience in building and operating distributed production storage systems at scale.
• At least 7 years of expertise in Linux and Kubernetes storage fundamentals (NFS, CSI).
• A minimum of 1 year of experience in integrating or developing High Performance Storage solutions (VAST, Weka, DDN, PowerScale).
• Profound knowledge of Linux storage and networking, including kernel and NFS-client layers.
• Familiarity with Kubernetes storage, CSI, storage classes, and persistent volumes.
• Experience in infrastructure-as-code practices and GitOps methodologies.
• Proficiency with Terraform/OpenTofu and ArgoCD or Flux.
• Hands-on experience with bare-metal hardware is a significant advantage.
• Capability to operate in hybrid, edge, and air-gapped environments.
• Understanding of Cluster API (CAPI), K0rdent, Harbor, and PKI/TLS considerations.
• Experience in building monitoring, alerting, and observability frameworks for storage systems.
• Opportunities for professional development and training.
• Participation in conferences and working groups.
• Company outings, happy hours, hackathons, and tech talks.
• Competitive compensation package complemented by a robust benefits plan.
appsoluts GmbH
SYNCREON
New Charter Technologies
Peraton
Get handpicked remote jobs straight to your inbox weekly.