
Engineering Manager, SRE
Posted 12 hours ago

Posted 12 hours ago
This is a fully remote position, open to applicants anywhere in the world.
• Lead the Site Reliability Engineering team, functioning as a 60% individual contributor and 40% in a leadership capacity.
• Oversee the onboarding, feedback, performance evaluation, career progression, and hiring of direct reports.
• Facilitate the career advancement of team members and promote alignment with company objectives.
• Act as the representative for the team within the engineering department and to senior leadership.
• Establish SRE goals, set priorities, and define delivery commitments.
• Manage the support rotation and the on-call schedule.
• Determine technical direction and review the outputs of the team.
• Be responsible for Remote's core infrastructure, which includes Kubernetes, AWS, PostgreSQL, DNS, TLS, and CI infrastructure.
• Enhance the reliability practices encompassing SLOs, error budgets, incident response, and observability.
• Collaborate with the Security team on threats, patching, infrastructure controls, and compliance obligations.
• Administer relationships with platform vendors, including renewals and commercial discussions with support from the Director.
• Optimize the balance between operational load and project delivery while expanding the SLO framework.
• Proven experience in leading an SRE, infrastructure, or platform engineering team.
• Accountability for the growth, performance, and career development of reports.
• Ability to coach both technical skills and soft skills.
• Experience in addressing underperformance directly and in a timely manner.
• Background in hiring engineers.
• Practical experience in site reliability, DevOps, or cloud infrastructure engineering.
• Proficiency with Kubernetes in a production environment.
• Experience with AWS at a significant scale.
• Hands-on experience in building and scaling AI infrastructure.
• Strong foundation in observability practices and principles.
• Familiarity with Infrastructure as Code using Terraform.
• Experience with CI/CD systems such as GitLab CI, GitHub Actions, or Jenkins.
• Skills in Docker and shell scripting.
• Experience in managing a reliability practice, including incident response, on-call duties, SLOs, and error budgets.
• Understanding and experience in working within regulated environments.
• Exceptional ability to prioritize between operational responsibilities and project tasks.
• Proficient written communication skills in a fully distributed, asynchronous setting.
• Capacity to establish relationships across various teams.
• Submission of application and CV in English is required.
• Work from anywhere.
• Flexible paid time off.
• Flexible working hours (we operate asynchronously).
• 16 weeks of paid parental leave.
• Budget allocated for co-working spaces, learning, and wellness (including gym memberships).
• Mental health support services.
• Stock options.
• Home office budget and IT equipment provided.
• Emphasis on life-work balance and schedule flexibility.
• Access to employee resource groups (Women, Disability, Queer, Minorities in Tech).
• Workplace accommodations available during interviews and beyond.
Reap
InTandem
Loancrate
Gainwell Technologies
Get handpicked remote jobs straight to your inbox weekly.