
Software Engineer II – Platform Reliability Engineering
Posted 23 hours ago

Posted 23 hours ago
This is a fully remote position, open to applicants in United States.
• Ensure the resilience, performance, and security of the enterprise Cloud Platform
• Execute reliability engineering projects
• Eliminate manual tasks through automation
• Engage in incident triage and conduct root cause analysis
• Assist in resolving systemic problems without blame through postmortems and problem management
• Integrate reliability into platforms via automation, change/incident/problem management, and destructive testing
• Support Service Level Objectives for highly available paved-path solutions
• Collaborate with UX, engineering, and product management team members
• Develop secure, reliable, and scalable software solutions
• Document, review, and uphold quality and change-control standards
• Aid in crafting developer-ready, comprehensible, and testable user stories
• Write code and scripts to automate infrastructure, monitoring services, and testing scenarios
• Create code and scripts for conducting destructive resiliency testing
• Configure and adjust programs and commercial off-the-shelf solutions
• Develop dashboards, logging, alerting, and proactive response mechanisms
• Collaborate in agile processes and forge partnerships across diverse teams
• Must be eighteen years of age or older
• Must be legally authorized to work in the United States
• 2+ years of relevant work experience in a related engineering field (Systems, Software, Operational) or in a reliability engineering domain
• Proficient understanding of ITIL processes and support/maintenance of production systems, including Change, Incident, and Problem Management
• Experience utilizing AI tools to minimize MTTD, MTTM, and MTTR
• Proficient in BASH, Python, Golang, TypeScript, Java, YAML, JSON, and HCL
• Familiarity with Terraform and Ansible
• Experience with Google Cloud Platform or similar projects and services, including infrastructure, Compute, Developer Tools, Security, and Identity and Access Management
• Experience with Prometheus, Grafana, and OpenTelemetry
• Understanding of Kubernetes/GKE and contemporary microservice architectures
• Knowledge of Unix and Linux operating systems
• Exposure to security tools and frameworks, including Wiz
• Experience conducting destructive, performance, and failure-scenario tests
• Knowledge of modern debugging and root cause analysis techniques
• Familiarity with version control systems
• Understanding of SLOs and core SRE principles and practices
• Strong communication and collaboration skills, including operational status communications, real-time stakeholder reporting, and documentation
• Bachelor's degree or equivalent in a field relevant to the position
• Competitive salary
• Comprehensive health benefits
• Opportunities for professional growth and development
• Flexible working hours and remote work options
• Collaborative and inclusive company culture
Natera
The Home Depot
Assist World
Seneca Holdings
Get handpicked remote jobs straight to your inbox weekly.