Senior Site Reliability Engineer, AUS

Posted 18 hours ago

This is a fully remote position, open to applicants in Australia.

📋 Description

• Take ownership of production reliability for customer-facing platforms and data services across Azure, colocation, and edge Kubernetes environments.

• Provide support to the entire organization through a shared Site Reliability Engineering (SRE) function across radar network and weather intelligence operations.

• Establish and enhance Service Level Indicators (SLIs), Service Level Objectives (SLOs), alerting standards, and operational metrics.

• Create shared observability dashboards and alerting systems for the entire fleet.

• Design and implement automated recovery and self-healing mechanisms for production systems.

• Optimize cluster resources and costs by right-sizing workloads and nodes, and migrating workloads away from Azure.

• Coordinate responses to production incidents, including troubleshooting, mitigation, communication, and postmortem analysis.

• Identify and resolve complex issues across application services, Kubernetes infrastructure, storage, and distributed systems.

• Promote multi-replica and multi-cluster high availability through effective workload placement, scheduling, failover, traffic routing, and data replication.

• Manage and enhance self-managed Kubernetes across cloud-hosted, colocation, and edge clusters.

• Execute Kubernetes upgrades, patching, cluster health management, node management, and production change management.

• Enhance observability, autoscaling, ingress, distributed storage, resiliency, and operational maturity.

• Design and validate Kubernetes workloads to ensure resiliency, scalability, and operational efficiency.

• Collaborate with software engineering teams on production readiness, deployment safety, resiliency, and operational visibility.

• Maintain deployment pipelines, Helm charts, Kubernetes manifests, and infrastructure automation.

• Support metrics, logging, distributed tracing, dashboarding, and alerting platforms.

• Conduct performance engineering and capacity planning for peak weather-event demands.

• Facilitate blameless postmortems and oversee complete operational follow-up actions.

• Enhance disaster recovery, failover, and business continuity across cloud, colocation, and edge environments.

• Advocate for automation, toil reduction, game days, production readiness reviews, and reliability best practices.

• Act as a senior technical resource and mentor.

• Participate in distinct rotating weekday and weekend on-call schedules, approximately every five weeks.


⛳️ Requirements

• A bachelor's degree in computer science, software engineering, or a related field; equivalent professional experience will be considered.

• At least 7 years of experience in Site Reliability Engineering, DevOps, Production Engineering, Platform Engineering, or a similar infrastructure-focused role, with a minimum of 4 years in a position formally titled Site Reliability Engineer or one that includes explicit SLO/error-budget accountability.

• Extensive, hands-on experience operating native Kubernetes; managed distributions such as AKS and EKS are also acceptable.

• Proven experience optimizing Kubernetes clusters, including right-sizing workloads and node pools, resource management, and minimizing infrastructure costs.

• Proficient in building dashboards, metrics pipelines, and alerting systems.

• Skilled in designing and operating workloads for safe horizontal scaling across multiple replicas, addressing idempotency, concurrency, and state handling.

• Experience in designing or managing multi-cluster high-availability architectures, including failover strategies, traffic routing, and cross-cluster service deployment.

• Background in supporting customer-facing production systems with responsibilities for uptime, reliability, and incident response.

• Experience in diagnosing and resolving production incidents across application, platform, and Kubernetes infrastructure layers.

• Familiarity with operating Kubernetes in environments beyond strictly managed cloud, including bare-metal, colocation, edge, or hybrid infrastructures.

• Knowledge of Kubernetes tooling and ecosystem technologies such as Rancher, Helm, autoscaling frameworks, observability stacks, or distributed storage systems.

• Strong grasp of infrastructure automation and Infrastructure as Code using tools such as Terraform and Ansible.

• Experience with CI/CD and production deployment pipelines; GitHub Actions is utilized at Climavision.

• Proficient in monitoring, logging, and observability platforms such as DataDog, Prometheus, Grafana, Loki, OpenTelemetry, or similar technologies.

• Experience in operating distributed systems and microservice-based architectures in production.

• Working knowledge of Microsoft Azure infrastructure.

• Strong troubleshooting skills across infrastructure, application, and platform layers.

• Experienced in participating in a structured production on-call rotation supporting business-critical systems.

• Familiarity with Jira, Confluence, and Microsoft Entra.

• Excellent written and verbal communication skills, including incident documentation and postmortem authoring.

• Experience in fast-paced engineering environments such as start-ups or scale-ups.

• Any employment offer is contingent upon the successful completion of a background check to meet company standards.


🏝️ Benefits

• Benefits of a dynamic and growing organization.

• A challenging, hands-on role that will have a real impact on the business.

• Competitive compensation.

• Comprehensive benefits package.

• 401(k) Savings Plan.

• Medical/Dental/Vision Benefits.

• Health Savings Account (HSA) and Flexible Spending Account (FSA).

• Unlimited Paid Time-off.

• 11 Paid Holidays.

• Paid Parental Leave.

• Company Paid Short-term Disability (STD).

• Company Paid Long-term Disability (LTD).

• Company Paid Life Insurance.

• Rotating weekday and weekend on-call schedule with separate rotations.

People also viewed

Colonist17 hours ago

DevOps Engineer

PT flagPortugal OnlyFreelanceDevOps & Site Reliability Engineer (SRE)
ApplyView job
Yopeso17 hours ago

Reliability Engineer / DevOps – Database Platform

RO flagRomania OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Orion Innovation17 hours ago

DevOps

MX flagMexico OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Climavision17 hours ago

Senior Site Reliability Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$130k – $170k/year
ApplyView job
Parasail18 hours ago

Senior Site Reliability Engineer

EuropeFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
NICE18 hours ago

Cloud Operations Engineer

GB flagUnited Kingdom OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers