
Senior Storage Production Engineer – DGX Cloud
Posted 6 days ago

Posted 6 days ago
This is a fully remote position, open to applicants in Australia.
• Design, implement, and maintain large-scale storage clusters.
• Ensure high availability, scalability, and data integrity.
• Develop and sustain storage monitoring, logging, and alerting systems.
• Enhance storage architectures for AI/ML workloads, focusing on low-latency access, efficient caching, and high-throughput performance.
• Oversee the storage service lifecycle, from design and deployment to operation and ongoing optimization.
• Provide consulting for system builds, automation frameworks, capacity management, and launch reviews.
• Monitor availability, latency, and system health utilizing predictive analytics and AI-driven automation.
• Optimize storage through techniques such as compression, deduplication, tiering, and intelligent workload placement.
• Scale storage using AI/ML-driven automation, policy-based tiering, and dynamic data migration.
• Implement encryption, access controls, and auditing mechanisms.
• Engage in sustainable incident response and conduct blameless root cause analysis.
• Participate in an on-call rotation to support storage and production systems.
• Bachelor's degree or equivalent experience in Computer Science, Storage Systems, or a related technical field.
• Over 8 years of practical experience.
• Proven experience with distributed and high-performance storage solutions, including clustered and parallel file systems, distributed object storage, and enterprise-grade storage systems.
• In-depth understanding of block, file, and object storage technologies, including their scalability, reliability, and performance characteristics.
• Experience with protocols such as NFS, SMB, iSCSI, S3, Fibre Channel, RDMA, and NVMe over Fabrics.
• Expertise in algorithms, data structures, complexity analysis, software design, and automating the maintenance of large-scale Linux-based storage systems.
• Proficient in one or more programming languages including C/C++, Java, Python, Go, NodeJS, and Bash.
• Hands-on experience with configuration management tools like Ansible, Chef, Puppet, and Terraform.
• Familiarity with InfluxDB, Prometheus, Grafana, and the Elastic stack.
• Excellent written and verbal communication skills.
• Strong work ethic, ability to work collaboratively, commitment to quality, and task completion.
• Knowledge of distributed storage systems, including replication strategies, erasure coding, capacity planning, performance tuning, and troubleshooting.
• Experience with version control systems such as Git, code review processes, pipelines, and CI/CD practices.
• Familiarity with Kubernetes, OpenStack, or hybrid cloud storage architectures.
• Capability to design automated storage migration, backup, and disaster recovery strategies.
• Full-time employment.
• Remote work arrangement in Australia.
Redox
TMS
Get handpicked remote jobs straight to your inbox weekly.