
Senior Storage Engineer
Posted Jul 27

Posted Jul 27
This is a fully remote position, open to applicants in Texas.
• Design, develop, and implement high-performance storage systems tailored for large GPU compute clusters, ensuring the storage architecture fulfills the throughput and latency requirements essential for AI training and inference tasks.
• Take ownership of the entire operational lifecycle of storage infrastructure, from initial deployment to daily performance management, capacity expansion, and end-of-life planning.
• Oversee and maintain storage networking fabrics that connect storage systems to GPU compute, ensuring that the interconnect remains a non-issue.
• Assess new storage hardware, software, and methodologies as the technology landscape evolves and IREN's footprint expands.
• Continuously monitor the health, performance, and capacity of storage systems; establish meaningful thresholds and proactively address issues before they escalate into incidents.
• Optimize storage systems for the unique I/O patterns associated with AI workloads and document effective strategies.
• Proactively plan and manage storage capacity, collaborating with teams to foresee demand ahead of new cluster deployments.
• Ensure high availability across storage tiers; design and validate resilience mechanisms to identify and mitigate single points of failure.
• Urgently diagnose and resolve storage incidents.
• Conduct comprehensive root cause analyses following significant incidents; create clear reports and implement changes to prevent future occurrences.
• Participate in on-call rotations to address storage-related production issues; serve as the escalation point for storage during critical incidents.
• Collaborate closely with technology and operations teams on new cluster builds, ensuring that storage is prepared and validated.
• Create and maintain technical documentation that is both accurate and up-to-date.
• Contribute to the technical assessment of storage vendors and products; clearly represent IREN's requirements in discussions with vendors and partners.
• A minimum of seven years of hands-on storage engineering experience in production environments, with substantial time spent managing storage at a significant scale.
• Proven experience in deploying, operating, and troubleshooting storage systems that cater to demanding compute workloads.
• Familiarity with data center environments; comfortable working in both the physical and software layers.
• Previous exposure to the specific requirements of AI training and inference I/O patterns is highly advantageous.
• Strong understanding of cloud platforms.
• In-depth practical knowledge of at least one major parallel or distributed file system utilized in high-performance environments.
• Excellent written and verbal communication skills, with the ability to effectively engage with executive leadership, service providers, and potentially community members.
• Exceptional time management and multitasking abilities, coupled with a strong sense of urgency.
• Demonstrated active listening and speaking skills.
• Post-secondary education in Computer Science or a related field.
• The Total Compensation package may include an annual incentive bonus and equity (long-term incentive).
• Relocation or living-out allowance / per diem (as applicable and based on the successful candidate's circumstances).
• 100% company-covered health insurance premiums (medical, dental, and vision) for employees, with 75% coverage for dependents.
• Company-paid short-term and long-term disability insurance.
• Availability of voluntary life, critical illness, and accident coverage.
• Health Savings Accounts (HSA) – available when combined with the High Deductible Health Plan.
• Access to the Employee Assistance Program and wellness resources.
• 401(k) retirement plan with company matching.
• Access to financial planning and legal services.
• Paid Time Off (PTO) and paid holidays.
• Opportunities for internal skills training and career advancement pathways.
• Professional development support for certifications, continuing education, or role-related training.
• Participation in company events and team-building activities.
Sigma Software Group
Plain Concepts
GitLab
Get handpicked remote jobs straight to your inbox weekly.