Remotery

Software Engineering Manager – Reliability Engineering, Store Systems

Posted Aug 5

This is a fully remote position, open to applicants in California, +6 more states.

📋 Description

• Ensure the resilience, performance, and security of Store Systems and associated applications.

• Lead incident triage, root cause analysis, and conduct blameless postmortems.

• Drive solutions to systemic problems to prevent recurrence.

• Engineer reliability into platforms via automation, change management, incident management, problem management, and destructive testing.

• Establish and enforce Service Level Objectives (SLOs), Service Level Indicators (SLIs), and Service Level Agreements (SLAs) for customer-facing workloads with high availability.

• Develop an infrastructure strategy that aligns with product goals, dependencies, and end-user requirements.

• Oversee infrastructure configuration, debugging, support, technology roll-outs, and the establishment of system software/hardware.

• Create and refine specifications for intricate technology solutions.

• Report on Systems Engineering progress to leadership.

• Manage relationships with vendors and handle hardware/software purchase requests.

• Prioritize escalations and requests from product teams and stakeholders.

• Produce internal solution documentation and proactively monitor systems for issues.

• Lead, mentor, coach, recruit, retain, and develop Systems Engineering professionals.

• Conduct performance reviews and manage individual development plans.

• Promote collaboration and eliminate obstacles.

• Advocate for the needs of end-users and stakeholders.

• Define and implement the long-term reliability roadmap.

• Oversee budgets, cloud expenditures, vendor negotiations, and resource allocation.

• Lead responses to high-severity incidents and implement preventative actions.

• Plan for capacity, champion chaos engineering, and support disaster recovery planning.

• Collaborate with software development, QA, security, and InfoSec teams to infuse reliability and security into the Software Development Life Cycle (SDLC).


⛳️ Requirements

• Must be at least eighteen years of age.

• Must have legal authorization to work in the United States.

• Bachelor’s degree or equivalent in a related field, or comparable experience.

• A minimum of 5 years of professional experience.

• Technical leadership experience in guiding and mentoring Site Reliability Engineering (SRE), infrastructure, and operations teams.

• Proven experience in defining and executing reliability roadmaps.

• Experience in recruiting, retaining, developing, and evaluating engineering talent.

• Experience managing budgets, cloud expenditures, vendor negotiations, and resource allocation.

• Proficiency in defining and enforcing SLOs, SLIs, and SLAs.

• Experience leading high-severity incidents, conducting blameless postmortems, and implementing preventative measures.

• Experience in capacity planning, chaos engineering, and disaster recovery.

• In-depth knowledge of Google Cloud Platform, AWS, or Microsoft Azure.

• Advanced skills in Terraform, Ansible, Chef, or Puppet.

• Proficient with monitoring, logging, and tracing tools such as Datadog, Prometheus, Grafana, Splunk, New Relic, or ELK.

• Experience overseeing Continuous Integration/Continuous Deployment (CI/CD) pipelines using Jenkins, GitLab CI, or GitHub Actions.

• Proficient in one or more programming languages such as Python, Go, Java, or Bash.

• Experience collaborating with software development, QA, security, and InfoSec teams.

• Ability to articulate technical metrics and incidents to executive leadership.

• Experience managing third-party SaaS and infrastructure providers.

• Familiarity with PCI-DSS, SOC2, HIPAA, and corporate security protocols.

• Understanding of least-privilege access models and audit logging.


🏝️ Benefits

• Remote/Virtual work arrangement.

• Overnight travel usually required only 5% to 20% of the time.

• Opportunities for mentoring, coaching, and professional development through learning activities and communities of practice.

• Support for career development, including individual development plans and clearly defined career paths.

• Performance feedback provided through annual and mid-year reviews.

People also viewed

Level Data5 hours ago

Senior DevOps Engineer

US flagMassachusetts OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$120k – $135k/year
ApplyView job
Level Data5 hours ago

DevOps Engineer II

US flagMassachusetts OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$95k – $110k/year
ApplyView job
Level Data5 hours ago

DevOps Engineer II

US flagMassachusetts OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$95k – $110k/year
ApplyView job
Level Data5 hours ago

Senior DevOps Engineer

US flagMassachusetts OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$120k – $135k/year
ApplyView job
Identity Digital Inc.9 hours ago

Staff DevOps Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$175k – $220k/year
ApplyView job
Identity Digital Inc.9 hours ago

Staff DevOps Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$175k – $220k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers