
Software Engineering Manager – Reliability Engineering, Store Systems
Posted Aug 5

Posted Aug 5
This is a fully remote position, open to applicants in California, +6 more states.
• Ensure the resilience, performance, and security of Store Systems and associated applications.
• Lead incident triage, root cause analysis, and conduct blameless postmortems.
• Drive solutions to systemic problems to prevent recurrence.
• Engineer reliability into platforms via automation, change management, incident management, problem management, and destructive testing.
• Establish and enforce Service Level Objectives (SLOs), Service Level Indicators (SLIs), and Service Level Agreements (SLAs) for customer-facing workloads with high availability.
• Develop an infrastructure strategy that aligns with product goals, dependencies, and end-user requirements.
• Oversee infrastructure configuration, debugging, support, technology roll-outs, and the establishment of system software/hardware.
• Create and refine specifications for intricate technology solutions.
• Report on Systems Engineering progress to leadership.
• Manage relationships with vendors and handle hardware/software purchase requests.
• Prioritize escalations and requests from product teams and stakeholders.
• Produce internal solution documentation and proactively monitor systems for issues.
• Lead, mentor, coach, recruit, retain, and develop Systems Engineering professionals.
• Conduct performance reviews and manage individual development plans.
• Promote collaboration and eliminate obstacles.
• Advocate for the needs of end-users and stakeholders.
• Define and implement the long-term reliability roadmap.
• Oversee budgets, cloud expenditures, vendor negotiations, and resource allocation.
• Lead responses to high-severity incidents and implement preventative actions.
• Plan for capacity, champion chaos engineering, and support disaster recovery planning.
• Collaborate with software development, QA, security, and InfoSec teams to infuse reliability and security into the Software Development Life Cycle (SDLC).
• Must be at least eighteen years of age.
• Must have legal authorization to work in the United States.
• Bachelor’s degree or equivalent in a related field, or comparable experience.
• A minimum of 5 years of professional experience.
• Technical leadership experience in guiding and mentoring Site Reliability Engineering (SRE), infrastructure, and operations teams.
• Proven experience in defining and executing reliability roadmaps.
• Experience in recruiting, retaining, developing, and evaluating engineering talent.
• Experience managing budgets, cloud expenditures, vendor negotiations, and resource allocation.
• Proficiency in defining and enforcing SLOs, SLIs, and SLAs.
• Experience leading high-severity incidents, conducting blameless postmortems, and implementing preventative measures.
• Experience in capacity planning, chaos engineering, and disaster recovery.
• In-depth knowledge of Google Cloud Platform, AWS, or Microsoft Azure.
• Advanced skills in Terraform, Ansible, Chef, or Puppet.
• Proficient with monitoring, logging, and tracing tools such as Datadog, Prometheus, Grafana, Splunk, New Relic, or ELK.
• Experience overseeing Continuous Integration/Continuous Deployment (CI/CD) pipelines using Jenkins, GitLab CI, or GitHub Actions.
• Proficient in one or more programming languages such as Python, Go, Java, or Bash.
• Experience collaborating with software development, QA, security, and InfoSec teams.
• Ability to articulate technical metrics and incidents to executive leadership.
• Experience managing third-party SaaS and infrastructure providers.
• Familiarity with PCI-DSS, SOC2, HIPAA, and corporate security protocols.
• Understanding of least-privilege access models and audit logging.
• Remote/Virtual work arrangement.
• Overnight travel usually required only 5% to 20% of the time.
• Opportunities for mentoring, coaching, and professional development through learning activities and communities of practice.
• Support for career development, including individual development plans and clearly defined career paths.
• Performance feedback provided through annual and mid-year reviews.
Level Data
Level Data
Level Data
Level Data
Get handpicked remote jobs straight to your inbox weekly.