
Member of Technical Staff – Storage Infrastructure
Posted 9 hours ago

Posted 9 hours ago
This is a fully remote position, open to applicants in California.
• Create and manage storage architectures for training datasets, checkpointing, inference artifacts, and collaborative research workflows.
• Implement and optimize parallel filesystems, object storage, and local NVMe caching to accommodate demanding AI workloads.
• Evaluate throughput, latency, metadata performance, and concurrent access using representative training and checkpoint workloads.
• Develop provisioning, capacity planning, lifecycle management, and operational automation for storage services.
• Design and validate replication, recovery, backup, and failure-handling protocols with defined durability and availability objectives.
• Identify performance and reliability challenges across applications, clients, networks, filesystems, and devices.
• Establish access controls, tenant separation, quotas, monitoring, and operational runbooks.
• Partner with compute and networking teams.
• Engage directly with customers who are pushing the limits of AI.
• Collaborate with the engineering team on systems that drive next-generation AI advancements.
• Over 3 years of experience in building or operating production distributed storage systems.
• Practical experience with at least one parallel or distributed filesystem or object storage platform, such as Lustre, BeeGFS, Ceph, or GPFS.
• Proficient in Linux administration and performance troubleshooting.
• Experience in automating infrastructure operations using Python, Go, Bash, or similar programming languages.
• Understanding of storage failure modes, data integrity, consistency, replication, and recovery procedures.
• Familiarity with block, file, and object storage semantics and the associated performance trade-offs.
• Experience with NVMe/SSD performance, filesystem tuning, I/O profiling, and benchmarking.
• Knowledge of high-throughput storage networking and the behavior of distributed clients.
• Experience in capacity forecasting, observability, alerting, and implementing safe maintenance practices.
• Understanding of authentication, authorization, encryption, and secure data lifecycle management.
• Experience in supporting large GPU training clusters and managing high-volume checkpoint workloads.
• Familiarity with S3-compatible object storage, data tiering, or distributed caching solutions.
• Experience with RDMA-enabled storage or GPUDirect Storage.
• Knowledge of Kubernetes storage integrations or SLURM environments.
• Expertise in storage cost optimization and contributions to open-source storage systems.
• Equity incentives
PerfectServe
Prime Intellect
Juniper Square
Get handpicked remote jobs straight to your inbox weekly.