
Team Leader, SRE
Posted 1 day ago

Posted 1 day ago
This is a fully remote position, open to applicants anywhere in the world.
• Guide the Site Reliability Engineering team, balancing 60% individual contributions with 40% leadership responsibilities.
• Manage the onboarding, feedback, performance evaluations, career growth, and recruitment of direct reports.
• Direct the team's alignment with company objectives and act as the representative for the team among engineering and senior leadership.
• Establish the technical direction and review the outputs of the team.
• Oversee SRE objectives, prioritization, support rotation, and the on-call structure.
• Supervise Kubernetes, AWS, PostgreSQL, DNS and TLS, as well as CI infrastructure.
• Cultivate the reliability practice, incorporating SLOs, error budgets, incident response, and observability.
• Collaborate with Security on threat management, patching, infrastructure controls, and compliance obligations.
• Handle platform vendor relationships, renewals, and commercial discussions.
• Enhance the balance between operational workload and project delivery, while expanding the SLO framework across teams.
• Proven experience in leading an SRE, infrastructure, or platform engineering team.
• Accountability for the growth, performance, and career development of team members.
• Proficient in coaching both technical skills and interpersonal abilities.
• Experience in addressing underperformance promptly and directly.
• Skilled in hiring engineers and evaluating engineering quality.
• Practical experience in site reliability, DevOps, or cloud infrastructure engineering.
• Proficient in Kubernetes within a production environment.
• Extensive experience with AWS at significant scale.
• Hands-on experience with AI development, enablement, and scaling AI infrastructure.
• Strong foundation in observability practices and principles.
• Familiarity with infrastructure as code using Terraform.
• Experience with CI/CD systems such as GitLab CI, GitHub Actions, or Jenkins.
• Knowledge of Docker and shell scripting.
• Background in managing a reliability practice, including incident response, on-call duties, SLOs, and error budgets.
• Understanding of and experience in regulated environments.
• Capability to balance operational tasks with project work.
• Excellent written communication skills for asynchronous collaboration.
• Ability to foster relationships across various teams.
• Submission of application materials in English is required.
• Work from any location.
• Flexible paid time off.
• Adaptable working hours (we operate asynchronously).
• 16 weeks of paid parental leave.
• Budget allocated for co-working spaces, professional development, and wellness (including gym memberships).
• Access to mental health support services.
• Stock options available.
• Home office budget and IT equipment provided.
• Emphasis on life-work balance and schedule flexibility.
• Participation in employee resource groups (Women, Disability, Queer, Minorities in Tech).
• Support for accommodations during the interview process and beyond.
VALCE Talent Solutions
4Pharma Ltd
MTP Brasil
BlackSky
Get handpicked remote jobs straight to your inbox weekly.