
Senior Site Reliability Engineer
Posted 4 days ago

Posted 4 days ago
This is a fully remote position, open to applicants in Australia.
• Drive the transformation of Upsun’s cloud application platform from conventional cloud operations to a proactive, automation-centric SRE model.
• Take ownership of vital engineering workstreams aimed at enhancing reliability, scalability, and operational efficiency across multi-cloud setups.
• Collaborate with engineering, product, and platform teams to integrate reliability and performance into the software delivery lifecycle.
• Identify architectural bottlenecks and promote infrastructure-as-code methodologies.
• Set up observability standards to ensure long-term system stability and uptime.
• Design monitoring, alerting, and logging systems utilizing Prometheus, Grafana, and ELK Stack.
• Create actionable SLIs/SLOs that align with key business metrics.
• Develop and implement robust automated infrastructure and workflows using Terraform and Ansible across AWS, GCP, and Azure.
• Enhance CI/CD pipeline architectures for rapid, secure, zero-downtime releases.
• Lead high-priority incident triage and facilitate blameless post-mortems.
• Apply preventative measures to strengthen system resilience.
• Collaborate with product and software engineering teams to integrate SRE practices into product roadmaps.
• Detect performance bottlenecks and assess technologies like eBPF and container orchestration.
• Participate in a four-week rotation balancing engineering and operations through hands-on troubleshooting and engineering innovation.
• Be on-call one week every 4–5 weeks, from 02:00–10:00 UTC, which includes a weekend shift.
• A minimum of 5 years of experience in Site Reliability Engineering, Cloud Operations, or DevOps.
• Demonstrated experience in ensuring reliability for production platforms at scale.
• Strong expertise in Go or Python for developing custom automation tools, custom controllers, or SRE platform components.
• Advanced hands-on knowledge of Linux operating system internals, kernel parameters, networking protocols, performance profiling, and system troubleshooting.
• In-depth expertise with AWS, GCP, Azure, or OpenStack.
• Experience with custom tools developed around cloud SDKs.
• Familiarity with declarative infrastructure tools like Terraform.
• Proven capability to foresee operational risks, make architectural trade-offs, and lead technical infrastructure initiatives with minimal supervision.
• Exceptional cross-functional communication skills and a history of building alignment while promoting an inclusive engineering culture.
• Must be legally authorized to work in Western Australia; visa sponsorship is not available.
• A successful background check is required.
• Bonus: experience with custom-built orchestration, edge, storage, and operational tools.
• Bonus: experience with Docker and managing production Kubernetes clusters or containerized deployment architectures.
• Bonus: familiarity with PaaS architectures or developer-facing cloud platforms.
• Flexible PTO.
• Company stock options.
• Professional development budget.
• Office equipment budget.
• Wellness budget.
• Annual team gatherings.
• Internet reimbursement.
• Inclusive parental leave.
• Remote work travel program.
• Flexible, open, and inclusive work environment.
• Accommodations available during the hiring process.
Koniag Government Services
FP Markets (First Prudential Markets)
Modern Campus
InRule
Get handpicked remote jobs straight to your inbox weekly.