
Software Engineer II – Reliability Engineering Tooling
Posted Sep 2

Posted Sep 2
This is a fully remote position, open to applicants in United States.
• Develop and maintain internal tools that support Developers, SREs, and Operations teams.
• Oversee the entire application lifecycle, encompassing development, testing, deployment, and operations.
• Serve as an SRE for applications you manage, treating internal associates as clients.
• Automate manual tasks and evaluate site health through relevant data analysis.
• Manage SOPs utilized for deploying and updating applications in GCP.
• Employ AI agents and capabilities to enhance change management practices and minimize incidents.
• Design, launch, and maintain production applications.
• Collaborate and work closely with UX, engineering, product management, and other product team members.
• Document, assess, and ensure compliance with quality and change control standards.
• Develop developer-friendly, comprehensible, and testable user stories in conjunction with the Product Team.
• Write code and scripts to automate infrastructure, monitoring services, test scenarios, and destructive testing.
• Configure and adapt programs and commercial off-the-shelf solutions as needed.
• Build dashboards, logging systems, alerts, and proactive response mechanisms.
• Participate in agile processes and enhance team productivity.
• Candidates must be at least eighteen years old.
• Must have legal authorization to work in the United States.
• Bachelor's degree or equivalent qualification in a relevant field of study.
• A minimum of 2 years of professional experience is required.
• 1-3 years of pertinent work experience in a related engineering or reliability engineering domain.
• Familiarity with ITIL processes and the support and maintenance of production systems, including Change, Incident, and Problem Management.
• Experience in leveraging AI technologies, including prompt engineering and creating custom AI agents and skills.
• Proficiency in Golang, JavaScript/TypeScript, and BASH.
• Experience in writing queries for relational or noSQL databases.
• Knowledge of infrastructure automation, CI/CD processes, Terraform, and GitHub Actions.
• Experience working within Google Cloud Platform or similar projects and services.
• Understanding of Kubernetes/GKE and contemporary microservice architectures.
• Familiarity with Prometheus, Grafana, and OpenTelemetry.
• Exposure to security tools such as Wiz and familiarity with security frameworks.
• Experience conducting destructive, performance, and failure-scenario tests.
• Proficiency in modern debugging and root cause analysis strategies.
• Experience utilizing version control systems.
• Understanding of SLOs along with fundamental SRE principles and practices.
• Strong communication and collaboration skills, including operational status reporting, real-time stakeholder updates, and documentation.
• Remote/Virtual work arrangement
• No travel required
Adapty.io
Oddball
General Dynamics Information Technology
StarTekk
Get handpicked remote jobs straight to your inbox weekly.