
Senior Site Reliability Engineer, AUS
Posted 18 hours ago

Posted 18 hours ago
This is a fully remote position, open to applicants in Australia.
• Take ownership of production reliability for customer-facing platforms and data services across Azure, colocation, and edge Kubernetes environments.
• Provide support to the entire organization through a shared Site Reliability Engineering (SRE) function across radar network and weather intelligence operations.
• Establish and enhance Service Level Indicators (SLIs), Service Level Objectives (SLOs), alerting standards, and operational metrics.
• Create shared observability dashboards and alerting systems for the entire fleet.
• Design and implement automated recovery and self-healing mechanisms for production systems.
• Optimize cluster resources and costs by right-sizing workloads and nodes, and migrating workloads away from Azure.
• Coordinate responses to production incidents, including troubleshooting, mitigation, communication, and postmortem analysis.
• Identify and resolve complex issues across application services, Kubernetes infrastructure, storage, and distributed systems.
• Promote multi-replica and multi-cluster high availability through effective workload placement, scheduling, failover, traffic routing, and data replication.
• Manage and enhance self-managed Kubernetes across cloud-hosted, colocation, and edge clusters.
• Execute Kubernetes upgrades, patching, cluster health management, node management, and production change management.
• Enhance observability, autoscaling, ingress, distributed storage, resiliency, and operational maturity.
• Design and validate Kubernetes workloads to ensure resiliency, scalability, and operational efficiency.
• Collaborate with software engineering teams on production readiness, deployment safety, resiliency, and operational visibility.
• Maintain deployment pipelines, Helm charts, Kubernetes manifests, and infrastructure automation.
• Support metrics, logging, distributed tracing, dashboarding, and alerting platforms.
• Conduct performance engineering and capacity planning for peak weather-event demands.
• Facilitate blameless postmortems and oversee complete operational follow-up actions.
• Enhance disaster recovery, failover, and business continuity across cloud, colocation, and edge environments.
• Advocate for automation, toil reduction, game days, production readiness reviews, and reliability best practices.
• Act as a senior technical resource and mentor.
• Participate in distinct rotating weekday and weekend on-call schedules, approximately every five weeks.
• A bachelor's degree in computer science, software engineering, or a related field; equivalent professional experience will be considered.
• At least 7 years of experience in Site Reliability Engineering, DevOps, Production Engineering, Platform Engineering, or a similar infrastructure-focused role, with a minimum of 4 years in a position formally titled Site Reliability Engineer or one that includes explicit SLO/error-budget accountability.
• Extensive, hands-on experience operating native Kubernetes; managed distributions such as AKS and EKS are also acceptable.
• Proven experience optimizing Kubernetes clusters, including right-sizing workloads and node pools, resource management, and minimizing infrastructure costs.
• Proficient in building dashboards, metrics pipelines, and alerting systems.
• Skilled in designing and operating workloads for safe horizontal scaling across multiple replicas, addressing idempotency, concurrency, and state handling.
• Experience in designing or managing multi-cluster high-availability architectures, including failover strategies, traffic routing, and cross-cluster service deployment.
• Background in supporting customer-facing production systems with responsibilities for uptime, reliability, and incident response.
• Experience in diagnosing and resolving production incidents across application, platform, and Kubernetes infrastructure layers.
• Familiarity with operating Kubernetes in environments beyond strictly managed cloud, including bare-metal, colocation, edge, or hybrid infrastructures.
• Knowledge of Kubernetes tooling and ecosystem technologies such as Rancher, Helm, autoscaling frameworks, observability stacks, or distributed storage systems.
• Strong grasp of infrastructure automation and Infrastructure as Code using tools such as Terraform and Ansible.
• Experience with CI/CD and production deployment pipelines; GitHub Actions is utilized at Climavision.
• Proficient in monitoring, logging, and observability platforms such as DataDog, Prometheus, Grafana, Loki, OpenTelemetry, or similar technologies.
• Experience in operating distributed systems and microservice-based architectures in production.
• Working knowledge of Microsoft Azure infrastructure.
• Strong troubleshooting skills across infrastructure, application, and platform layers.
• Experienced in participating in a structured production on-call rotation supporting business-critical systems.
• Familiarity with Jira, Confluence, and Microsoft Entra.
• Excellent written and verbal communication skills, including incident documentation and postmortem authoring.
• Experience in fast-paced engineering environments such as start-ups or scale-ups.
• Any employment offer is contingent upon the successful completion of a background check to meet company standards.
• Benefits of a dynamic and growing organization.
• A challenging, hands-on role that will have a real impact on the business.
• Competitive compensation.
• Comprehensive benefits package.
• 401(k) Savings Plan.
• Medical/Dental/Vision Benefits.
• Health Savings Account (HSA) and Flexible Spending Account (FSA).
• Unlimited Paid Time-off.
• 11 Paid Holidays.
• Paid Parental Leave.
• Company Paid Short-term Disability (STD).
• Company Paid Long-term Disability (LTD).
• Company Paid Life Insurance.
• Rotating weekday and weekend on-call schedule with separate rotations.
Colonist
Yopeso
Orion Innovation
Climavision
Get handpicked remote jobs straight to your inbox weekly.