
IT Service Reliability and Operational Excellence Manager
Posted Sep 8

Posted Sep 8
This is a fully remote position, open to applicants in New Zealand.
• Take charge of the reliability, resilience, and performance of live IT services, prioritizing customer and business outcomes.
• Spearhead Application Operations and Site Reliability Engineering, fostering enhanced service ownership, engineering accountability, and continuous improvement.
• Leverage AI, AIOps, and automation to anticipate issues sooner, minimize operational toil, and boost self-healing capabilities.
• Convert operational challenges into improvements in engineering, testing, environment, release, and platform processes.
• Transition the operating model from reactive, application-centric support to one focused on service and value-stream ownership.
• Integrate insights from live production into Quality, Release, NPE, Service Desk, Architecture, and Engineering for ongoing organizational learning and enhancement.
• Enhance observability and reliability engineering by utilizing service health metrics, SLIs, and SLOs.
• Oversee operational risk, service continuity, and disaster recovery initiatives.
• Revolutionize strategic partner performance towards automation, engineering advancements, and shared service outcomes.
• Elevate operational readiness, problem management, and resilience, ensuring that changes are implemented smoothly and that recurring failures are systematically eradicated.
• A tertiary qualification in Software Engineering, Computer Science, Information Technology, or a related field, or equivalent industry experience.
• 8–10 years of experience in IT operations, application support, reliability engineering, or service management within a large and complex technology environment.
• Demonstrated leadership of large-scale Application Operations and Site Reliability Engineering teams in a 24/7 service environment.
• Proven experience in managing onshore and offshore outsourced service providers through robust commercial governance and continuous service improvement.
• Strong practical expertise in SRE and observability, including SLIs, SLOs, error budgets, service mapping, monitoring, logging, tracing, alerting, capacity, and resilience engineering.
• Established experience in transforming operations from reactive support to proactive, automated, and value-stream-aligned service management.
• Proven leadership in managing major incidents, problem management, service transition, operational risk, service continuity, disaster recovery, and IT service management.
• Experience in applying AIOps, automation, and machine learning to minimize toil, identify anomalies, correlate events, and enhance remediation efforts.
• Strong influencing capabilities across Architecture, Engineering, Delivery, Quality Assurance, Environment Management, Cyber Security, business stakeholders, and strategic partners.
• Highly developed skills in people leadership, commercial management, communication, and data-driven problem-solving, with the ability to translate customer and business impacts into clear reliability priorities.
• Most roles provide the flexibility to work from home and adjust hours to accommodate work and family commitments.
• A fully subsidized Southern Cross health insurance plan for you and your family.
• Lifestyle leave options, allowing you to purchase an additional week or two of annual leave.
• Discounts on One New Zealand products, services, and much more.
• A Rainbow Tick certified and diversity-focused workplace.
Vitable Health
The Cigna Group
AmpiFire
Worldwide Clinical Trials
Get handpicked remote jobs straight to your inbox weekly.