
Senior Production Engineer
Posted Sep 4

Posted Sep 4
This is a fully remote position, open to applicants in Australia.
• Lead a team of engineers with a focus on Site Reliability Engineering (SRE) principles.
• Design, develop, and sustain complex systems on a large scale.
• Foster an engineering culture that prioritizes delivery while upholding high standards for stability, performance, security, and scalability.
• Oversee monitoring, alerting, investigation, and resolution of production infrastructure issues across public cloud platforms, including GCP and AWS.
• Discover opportunities and implement automated, scalable solutions to enhance code and operational tasks, minimizing toil.
• Act as a technical liaison during critical incidents.
• Promote cross-functional communication, monitor alert trends, and conduct root cause analysis (RCA) to facilitate structural enhancements.
• Implement production changes and security requests, which encompass code deployments and patching pipelines.
• Collaborate with Product Management and Engineering Management to ensure technical solutions and execution strategies are aligned with business objectives and roadmap priorities.
• Mentor, coach, and develop engineers throughout the organization.
• A minimum of 5 years of experience in software engineering, DevOps, or Site Reliability Engineering (SRE) focusing on building and managing production-ready configurations at internet scale.
• A solid background in system architecture, distributed microservices, and self-healing systems with effective resource and network utilization.
• Advanced knowledge of executing containerized workloads using Kubernetes orchestration and Docker in public cloud hosting environments.
• Proficient in structural and automated scripting utilizing object-oriented programming languages.
• Strong experience in Python and Go is preferred.
• Comprehensive understanding of the modern web serving stack and infrastructure automation tools, including Linux, Nginx/Apache, MySQL, PHP, Ansible, and Terraform.
• Hands-on experience with implementing metrics, monitoring, and alerting systems such as Prometheus or EFK stacks.
• Excellent analytical, troubleshooting, and root cause analysis skills.
• Capability to thrive in a team-oriented environment.
• Company Stock Options (Every employee is an owner in the company).
• Superannuation Program.
• Employee Assistance Program.
• Supplemental Maternity & Paternity Pay.
• Generous Vacation Time (Who doesn't like time off).
• One-time $745 AUS Home Office Stipend.
• Company Wellness Days.
NVIDIA
NVIDIA
Redox
NVIDIA
Get handpicked remote jobs straight to your inbox weekly.