
Engineering Team Leader – Site Reliability Engineering
Posted 2 days ago

Posted 2 days ago
This is a fully remote position, open to applicants in Poland.
• Cultivate and expand a high-performing Site Reliability Engineering team, promoting technical excellence, ownership, and continuous improvement.
• Establish and steer the SRE platform strategy in partnership with infrastructure and development teams.
• Supervise the organization-wide reliability strategy, ensuring architectural alignment, scalability, and the adoption of SRE best practices.
• Take ownership of 24/7 on-call and incident management processes.
• Manage incident tooling, develop operational procedures, ensure compliance with regulations, oversee reporting, and enhance incident response strategies.
• Define and monitor team performance KPIs and management metrics.
• Oversee the design, development, and advancement of the organization-wide observability ecosystem.
• Lead the implementation of standardized telemetry, including structured logging, distributed tracing, and intelligent sampling.
• Deconstruct complex projects into actionable tasks and promote incremental value delivery.
• Mentor team members and facilitate their professional development.
• Conduct workshops, foster a technical community, and mediate conflicts.
• Promote operational excellence and a culture of reliability, lead incident management efforts, and advocate for post-mortems.
• Manage technical debt and align team output with the organization's objectives.
• Several years of experience in Site Reliability Engineering, Infrastructure, or DevOps roles managing large-scale, distributed environments.
• Proven experience in formal management, leading, mentoring, and developing high-performing SRE or DevOps engineering teams.
• Extensive background in building and maintaining scalable, reliable, and observable infrastructure systems across Azure, Kubernetes, and on-premises environments.
• Demonstrated capability to deliver comprehensive reliability strategies, drive architectural enhancements, and manage large-scale technical projects from design to production.
• Proven ability to foster cultural change, act as a strategic partner to product engineering teams, and collaborate effectively with distributed/remote teams.
• Strong proficiency in Python for scalable automation, internal tools, and scripting purposes.
• Expertise in Kubernetes and Ansible for configuration management.
• Ability to design resilient infrastructure both on Azure and on-premises.
• Deep expertise in standardized telemetry systems.
• Mastery of tools such as Prometheus, Grafana, OTEL, ELK, Tempo, Thanos, or similar technologies.
• Experience utilizing AI/ML for AIOps, anomaly detection, log analysis, and optimizing reliability workflows.
• Nice to have: Experience with commercial APM platforms such as Datadog, Splunk, or New Relic, and chaos engineering tools.
• Nice to have: Proficiency in cloud cost management and FinOps principles.
• Significant influence on the development of the company and its products.
• Work alongside an experienced team eager to share knowledge.
• Clear development vision supported by regular feedback and well-defined career paths.
• Regular team-building events.
• A training budget for courses and conferences that interest you.
• An additional day off on your birthday.
• An extra day off for parents.
• Equipment customized to your requirements.
• Private medical care and group insurance.
• Access to an e-learning platform for English language learning and a benefits platform.
• Access to a wellbeing platform, including opportunities for workshops and private therapy sessions.
• Flexibility to work remotely, from the office in Warsaw, or from a coworking space in your city.
CWILL
a37
GT
Sigma Software Group
Get handpicked remote jobs straight to your inbox weekly.