AI Storage Infrastructure Engineer

Posted Sep 11

This is a fully remote position, open to applicants in United States.

📋 Description

• Integrate high-performance NFS-based storage solutions such as VAST and Dell PowerScale into Kubernetes clusters utilizing CSI, storage classes, and persistent volumes.

• Optimize NFS data paths, including mount options, nconnect/RDMA, Linux clients, and network settings, to support high-throughput, low-latency GPU/AI workloads.

• Deploy and manage storage services and operators effectively.

• Oversee storage capacity, quotas, snapshots, and their lifecycle management.

• Configure and enhance Linux systems for storage workloads, focusing on drivers, file-system layout, network tuning, and kernel parameters.

• Provide storage integration for k0s-based Kubernetes through Cluster API and K0rdent management/child-cluster architectures.

• Operate storage solutions in fully disconnected air-gapped environments, including managing Harbor artifact/mirror connectivity and PKI/TLS aspects.

• Automate storage provisioning and configuration using Terraform/OpenTofu and GitOps pipelines with ArgoCD or Flux.

• Develop monitoring, alerting, and observability systems for storage performance, capacity, and health.

• Troubleshoot and resolve performance, reliability, and scaling challenges throughout the storage stack.

• Establish operational standards and facilitate communication across teams.


⛳️ Requirements

• Over 7 years of experience in Site Reliability Engineering (SRE) or hardware/storage infrastructure operations.

• More than 5 years of experience in building and operating distributed production storage systems at scale.

• At least 7 years of expertise in Linux and Kubernetes storage fundamentals (NFS, CSI).

• A minimum of 1 year of experience in integrating or developing High Performance Storage solutions (VAST, Weka, DDN, PowerScale).

• Profound knowledge of Linux storage and networking, including kernel and NFS-client layers.

• Familiarity with Kubernetes storage, CSI, storage classes, and persistent volumes.

• Experience in infrastructure-as-code practices and GitOps methodologies.

• Proficiency with Terraform/OpenTofu and ArgoCD or Flux.

• Hands-on experience with bare-metal hardware is a significant advantage.

• Capability to operate in hybrid, edge, and air-gapped environments.

• Understanding of Cluster API (CAPI), K0rdent, Harbor, and PKI/TLS considerations.

• Experience in building monitoring, alerting, and observability frameworks for storage systems.


🏝️ Benefits

• Opportunities for professional development and training.

• Participation in conferences and working groups.

• Company outings, happy hours, hackathons, and tech talks.

• Competitive compensation package complemented by a robust benefits plan.

People also viewed

appsoluts GmbH1 day ago

Web, Backend & Infrastructure Developer

DE flagGermany OnlyFull-timeInfrastructure Engineer€50k – €75k/year
ApplyView job
SYNCREON1 day ago

Senior Mainframe Systems Programmer – Infrastructure Engineering

US flagNew Jersey OnlyFull-timeInfrastructure Engineer
ApplyView job
New Charter Technologies1 day ago

Infrastructure Specialist

US flagNorth Carolina OnlyFull-timeInfrastructure Engineer$76k/year
ApplyView job
Peraton1 day ago

Oracle Cloud Infrastructure Architect

US flagUnited States OnlyFull-timeInfrastructure Engineer$112k – $179k/year
ApplyView job
centrapay1 day ago

Infrastructure Engineer

NZ flagNew Zealand OnlyFull-timeInfrastructure EngineerNZ$90k – NZ$125k/year
ApplyView job
Credit Acceptance2 days ago

Staff Systems Engineer, Infrastructure

US flagMichigan OnlyFull-timeInfrastructure Engineer$122k – $178.9k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers