
Senior Software Engineer – Platform Reliability Engineering
Posted 23 hours ago

Posted 23 hours ago
This is a fully remote position, open to applicants in United States.
• Ensure the resilience, performance, and security of the enterprise Cloud Platform.
• Collaborate with partner Reliability Engineering and Enablement teams to support essential Engineering Experience services.
• Execute reliability engineering initiatives effectively.
• Eliminate toil through automation processes.
• Provide mentorship to junior engineers and facilitate technical discussions.
• Lead incident triage efforts and conduct root cause analyses.
• Drive permanent resolutions to systemic issues through blameless postmortems and effective problem management.
• Develop, test, deploy, and maintain software solutions.
• Create functional and destructive test suites to enable swift production deployment.
• Work collaboratively with team members using agile methodologies.
• Partner with Product Teams to ensure user stories are valuable, developer-ready, comprehensible, and testable.
• Design and implement destructive, performance, and failure-scenario tests, including team drills to validate operational readiness.
• Set and maintain Service Level Objectives for highly available paved-path solutions.
• Must be at least eighteen years old.
• Must have legal authorization to work in the United States.
• A minimum of 3 years of professional experience is required.
• Bachelor's degree or an equivalent qualification in a related field is necessary.
• Over 4 years of relevant work experience in a related engineering field or reliability engineering domain is preferred.
• Familiarity with ITIL processes and production system support and maintenance, including Change, Incident, and Problem Management.
• Experience in utilizing AI tools to reduce MTTD, MTTM, and MTTR.
• Proficiency in BASH, Python, Golang, TypeScript, Java, YAML, JSON, and HCL.
• Experience with Terraform and Ansible.
• Proven track record of managing Google Cloud Platform or similar projects and services.
• Knowledge of Prometheus, Grafana, and OpenTelemetry.
• Strong understanding of Kubernetes/GKE and contemporary microservice architectures.
• Familiarity with Unix and Linux operating systems.
• Experience with Wiz and security frameworks.
• Capability in designing and executing destructive, performance, and failure-scenario tests.
• Proficiency in debugging and root cause analysis techniques.
• Familiarity with version control systems.
• Strong understanding of SLOs and core SRE principles and practices.
• Experience in producing operational status communications, real-time reporting, documentation, and peer mentorship.
• Remote/Virtual work arrangement.
• No travel is required.
• Opportunities for mentorship and professional development.
Natera
The Home Depot
Assist World
Seneca Holdings
Get handpicked remote jobs straight to your inbox weekly.