
Senior Manager, Storage Production Engineering
Posted Jul 29

Posted Jul 29
This is a fully remote position, open to applicants in California.
• Leading and mentoring a team of Storage Production Engineers to foster a collaborative, inclusive, and learning-driven environment.
• Designing, implementing, and enhancing large-scale storage systems, including distributed storage, parallel file systems, and object storage solutions.
• Utilizing automation, monitoring, and analytics to enhance the reliability, efficiency, and operability of storage services.
• Taking responsibility for capacity planning, data lifecycle management, cost awareness, and ensuring high availability and disaster recovery strategies for storage systems.
• Assessing and integrating modern storage technologies such as NVMe over Fabrics, RDMA, high-speed interconnects, and cloud-based storage solutions.
• Leading incident response efforts and conducting root cause analysis for storage-related issues, implementing changes to mitigate future occurrences.
• Collaborating with engineering, DevOps, and AI/ML teams to optimize data pipelines, access patterns, and overall workflow performance.
• Bachelor’s or Master’s degree in Computer Science, Storage Systems, or a related technical discipline, or equivalent professional experience.
• Over 12 years of experience in large-scale storage architecture, operations, production engineering, or infrastructure roles.
• At least 6 years of experience in people management or technical leadership within storage, infrastructure, or site reliability teams.
• Direct experience in managing infrastructure operations, including on-call responsibilities, incident response, ongoing maintenance, troubleshooting, and optimization of production systems, alongside managing SLOs and operational KPIs.
• Hands-on experience with parallel file systems (such as Lustre or GPFS), distributed storage (e.g., Ceph or MinIO), and enterprise object or NAS platforms (such as S3-compatible systems, NetApp, or Pure Storage).
• Strong understanding of block, file, and object storage, including performance tuning, data protection, and designing for high availability.
• Familiarity with storage networking and protocols such as NFS, SMB, iSCSI, Fibre Channel, RDMA, and NVMe-oF.
• Practical experience with automation and infrastructure as code using tools like Terraform, Ansible, or Puppet.
• Solid knowledge of monitoring and observability tools (e.g., Prometheus, InfluxDB, or Elastic stack), logging, and alerting mechanisms used for operating and enhancing storage systems.
• Medical, dental, and vision coverage.
• Mental health resources.
• Retirement and 401(k) plans.
• Employee stock purchase plan.
• Paid time off and holidays.
• Family and caregiving leave.
• A variety of wellness and development programs.
Vultr
Canva
Canva
NVIDIA
Get handpicked remote jobs straight to your inbox weekly.