Member of Technical Staff – Storage Infrastructure

atPrime IntellectRemoteUS flagCaliforniaFull-timeSoftware EngineerLead$150k – $300k/year

Posted 9 hours ago

This is a fully remote position, open to applicants in California.

📋 Description

• Create and manage storage architectures for training datasets, checkpointing, inference artifacts, and collaborative research workflows.

• Implement and optimize parallel filesystems, object storage, and local NVMe caching to accommodate demanding AI workloads.

• Evaluate throughput, latency, metadata performance, and concurrent access using representative training and checkpoint workloads.

• Develop provisioning, capacity planning, lifecycle management, and operational automation for storage services.

• Design and validate replication, recovery, backup, and failure-handling protocols with defined durability and availability objectives.

• Identify performance and reliability challenges across applications, clients, networks, filesystems, and devices.

• Establish access controls, tenant separation, quotas, monitoring, and operational runbooks.

• Partner with compute and networking teams.

• Engage directly with customers who are pushing the limits of AI.

• Collaborate with the engineering team on systems that drive next-generation AI advancements.


⛳️ Requirements

• Over 3 years of experience in building or operating production distributed storage systems.

• Practical experience with at least one parallel or distributed filesystem or object storage platform, such as Lustre, BeeGFS, Ceph, or GPFS.

• Proficient in Linux administration and performance troubleshooting.

• Experience in automating infrastructure operations using Python, Go, Bash, or similar programming languages.

• Understanding of storage failure modes, data integrity, consistency, replication, and recovery procedures.

• Familiarity with block, file, and object storage semantics and the associated performance trade-offs.

• Experience with NVMe/SSD performance, filesystem tuning, I/O profiling, and benchmarking.

• Knowledge of high-throughput storage networking and the behavior of distributed clients.

• Experience in capacity forecasting, observability, alerting, and implementing safe maintenance practices.

• Understanding of authentication, authorization, encryption, and secure data lifecycle management.

• Experience in supporting large GPU training clusters and managing high-volume checkpoint workloads.

• Familiarity with S3-compatible object storage, data tiering, or distributed caching solutions.

• Experience with RDMA-enabled storage or GPUDirect Storage.

• Knowledge of Kubernetes storage integrations or SLURM environments.

• Expertise in storage cost optimization and contributions to open-source storage systems.


🏝️ Benefits

• Equity incentives

People also viewed

PerfectServe9 hours ago

Mobile Developer I

US flagUnited States OnlyFull-timeSoftware Engineer$75k – $100k/year
ApplyView job
Prime Intellect9 hours ago

Member of Technical Staff – Datacenter Networking

US flagCalifornia OnlyFull-timeSoftware Engineer$150k – $300k/year
ApplyView job
Juniper Square9 hours ago

Director, Engineering – Platform

US flagUnited States, +1 more countryFull-timeSoftware Engineer$230k – $285k/year
ApplyView job
Certsys9 hours ago

Oracle APEX Developer

BR flagBrazil OnlyFull-timeSoftware Engineer
ApplyView job
Alterdata Software9 hours ago

Mid-Level Delphi Developer

BR flagBrazil OnlyFull-timeSoftware Engineer
ApplyView job
Valiant Solutions9 hours ago

Software Engineer

US flagUnited States OnlyFull-timeSoftware Engineer
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers