
Head of SRE
Posted Jul 28

Posted Jul 28
This is a fully remote position, open to applicants in Europe.
• Take ownership and spearhead all strategies, standards, and execution related to Site Reliability Engineering (SRE).
• Foster a culture of SRE and operational excellence throughout engineering teams.
• Assess the existing infrastructure and operational framework; redesign and reconstruct as necessary.
• Design, deploy, and sustain scalable, secure production environments.
• Establish and implement Service Level Indicators (SLIs), Service Level Objectives (SLOs), and uptime goals.
• Create strong monitoring, alerting, and observability practices.
• Develop and implement incident management, root cause analysis (RCA), and postmortem processes.
• Construct and oversee sustainable on-call frameworks and escalation procedures.
• Automate the software delivery lifecycle to enhance release predictability and safety.
• Generate reproducible environments and Infrastructure as Code (IaaC) provisioning templates.
• Enhance system performance, availability, and reliability.
• Support and operationalize data platforms and machine learning (ML) workloads.
• Collaborate closely with QA and Engineering leadership to elevate release quality and stability.
• Ensure infrastructure complies with enterprise-level security and regulatory standards.
• Recruit, manage, and mentor a team of SRE engineers.
• Demonstrated hands-on experience in Site Reliability Engineering, Production Engineering, or a comparable role.
• Strong practical knowledge of cloud infrastructure (preferably AWS or Azure), Infrastructure as Code (Terraform), and Kubernetes.
• Experience in establishing or enhancing SRE practices within an organization.
• Proven track record of improving uptime, reliability, and operational workflows.
• Comprehensive understanding of Continuous Integration/Continuous Deployment (CI/CD), development experience, infrastructure-as-code, and automation.
• Experience in designing on-call processes and incident response frameworks.
• Managed at least one team of SRE engineers.
• Excellent communication skills, capable of influencing across various teams.
• Experience in supporting data platforms and ML systems in production settings.
• MLOps experience, including model deployment, monitoring, and retraining workflows.
• Lead by example, act swiftly, and make informed, data-driven decisions.
• Continuous drive for improvement, always focused on delivering tangible value to customers.
CWILL
a37
GT
Sigma Software Group
Get handpicked remote jobs straight to your inbox weekly.