
AI Storage Infrastructure Engineer
Posted Sep 11

Posted Sep 11
This is a fully remote position, open to applicants in United States.
• Deploy, integrate, and manage high-performance storage solutions for GPU-accelerated computing and AI platforms.
• Incorporate NFS-based high-performance storage solutions like VAST and Dell PowerScale into Kubernetes clusters utilizing CSI, storage classes, and persistent volumes.
• Optimize NFS data paths, including mount options, nconnect/RDMA, Linux clients, and network configurations.
• Deploy and oversee storage services and operators.
• Manage storage capacity, quotas, snapshots, and lifecycle.
• Configure and enhance Linux systems for storage workloads, covering drivers, file systems, network configurations, and kernel parameters.
• Deliver storage integration for k0s-based Kubernetes utilizing Cluster API and K0rdent management/child cluster architectures.
• Operate storage in completely disconnected air-gapped environments, considering Harbor artifact/mirror connectivity and PKI/TLS factors.
• Automate storage provisioning and configuration using Terraform/OpenTofu and GitOps pipelines with ArgoCD or Flux.
• Develop monitoring, alerting, and observability systems for storage performance, capacity, and health.
• Diagnose and resolve performance, reliability, and scalability challenges across the storage stack.
• Establish operational standards and collaborate across teams.
• Over 7 years of experience in Site Reliability Engineering (SRE) or hardware/storage infrastructure operations.
• More than 5 years of experience in building and operating large-scale distributed storage systems.
• At least 7 years of experience with Linux and Kubernetes storage principles (NFS, CSI).
• A minimum of 1 year of experience in integrating or developing High-Performance Storage solutions (VAST, Weka, DDN, PowerScale).
• Extensive knowledge of Linux storage and networking fundamentals, down to the kernel and NFS-client layers.
• Proficiency in Kubernetes storage.
• Experience with infrastructure-as-code and GitOps methodologies.
• Proven ability to diagnose performance and reliability issues from end to end.
• Experience with bare-metal hardware is highly advantageous.
• Strong communication skills across teams.
• Opportunities for professional development and training.
• Participation in conferences and working groups.
• Company outings, happy hours, hackathons, and tech talks.
• Competitive compensation package coupled with a robust benefits plan.
• Collaborate with exceptionally passionate, talented, and engaging colleagues.
• Be part of pioneering open-source innovation.
• High-energy environment that values openness, collaboration, risk-taking, and continuous growth.
appsoluts GmbH
SYNCREON
New Charter Technologies
Peraton
Get handpicked remote jobs straight to your inbox weekly.