
Staff Software Engineer – Customer Reliability Engineering
Posted Sep 2

Posted Sep 2
This is a fully remote position, open to applicants in United States.
• Oversee the design and execution of foundational systems to ensure reliability, scalability, performance, and efficiency.
• Create enterprise blueprints for observability, automation, and cloud infrastructure architecture.
• Lead the selection, configuration, security, resilience, performance optimization, and monitoring of production tools.
• Establish Service Level Objectives, manage error budgets, and facilitate blameless post-incident review practices.
• Develop, test, deploy, and maintain software solutions.
• Create both functional and destructive test suites to facilitate swift production deployments.
• Approach technical challenges with a broad, global perspective.
• Collaborate with Product Teams to ensure user stories are ready for developers, clear, and testable.
• Engage in agile processes alongside team members.
• Address inquiries from product and engineering teams.
• Provide guidance to junior engineers on modern software development frameworks and lead technical discussions.
• Identify gaps within the team and propose enhancements for productivity.
• Influence technical decisions across teams without formal authority.
• Mentor engineers at all experience levels and raise engineering standards.
• Typically reports to a Software Engineering Manager or Senior Manager.
• Generally has no direct reports.
• Must be at least eighteen years of age.
• Must have legal authorization to work in the United States.
• Bachelor's degree or equivalent in a relevant field.
• Minimum of 3 years of professional experience.
• Preferred 8+ years of applicable experience in Cloud Operations, Site Reliability Engineering, DevOps, or Software Engineering within a high-scale, distributed setting.
• Extensive expertise in designing, deploying, and managing high-availability, multi-region production architectures on Google Cloud Platform or AWS/Azure.
• Proficient in Go, Python, or Java at a production level.
• In-depth knowledge of observability tools like Datadog, Prometheus, Grafana, or Splunk.
• Competence in defining and implementing SLIs, SLOs, and alerting strategies.
• Mastery of Infrastructure as Code tools such as Terraform or CloudFormation.
• Experience with CI/CD pipeline automation using GitHub Actions or Jenkins.
• Advanced operational expertise with Kubernetes, container orchestration, and microservices architectures.
• Proven experience in leading incident response, conducting root-cause analysis, and facilitating blameless post-mortems.
• Ability to steer the technical direction of intricate, cross-team initiatives.
• Established history of mentoring junior and mid-level engineers.
• Excellent communication skills and a strong ability to collaborate across Product, UX, Architecture, Security, and Engineering teams.
• No travel required.
• Comfortable indoor work environment.
• Frequent opportunities for movement during work.
Adapty.io
Oddball
General Dynamics Information Technology
StarTekk
Get handpicked remote jobs straight to your inbox weekly.