Remotery

Head of SRE

Posted Jul 28

This is a fully remote position, open to applicants in Europe.

📋 Description

• Take ownership and spearhead all strategies, standards, and execution related to Site Reliability Engineering (SRE).

• Foster a culture of SRE and operational excellence throughout engineering teams.

• Assess the existing infrastructure and operational framework; redesign and reconstruct as necessary.

• Design, deploy, and sustain scalable, secure production environments.

• Establish and implement Service Level Indicators (SLIs), Service Level Objectives (SLOs), and uptime goals.

• Create strong monitoring, alerting, and observability practices.

• Develop and implement incident management, root cause analysis (RCA), and postmortem processes.

• Construct and oversee sustainable on-call frameworks and escalation procedures.

• Automate the software delivery lifecycle to enhance release predictability and safety.

• Generate reproducible environments and Infrastructure as Code (IaaC) provisioning templates.

• Enhance system performance, availability, and reliability.

• Support and operationalize data platforms and machine learning (ML) workloads.

• Collaborate closely with QA and Engineering leadership to elevate release quality and stability.

• Ensure infrastructure complies with enterprise-level security and regulatory standards.

• Recruit, manage, and mentor a team of SRE engineers.


⛳️ Requirements

• Demonstrated hands-on experience in Site Reliability Engineering, Production Engineering, or a comparable role.

• Strong practical knowledge of cloud infrastructure (preferably AWS or Azure), Infrastructure as Code (Terraform), and Kubernetes.

• Experience in establishing or enhancing SRE practices within an organization.

• Proven track record of improving uptime, reliability, and operational workflows.

• Comprehensive understanding of Continuous Integration/Continuous Deployment (CI/CD), development experience, infrastructure-as-code, and automation.

• Experience in designing on-call processes and incident response frameworks.

• Managed at least one team of SRE engineers.

• Excellent communication skills, capable of influencing across various teams.

• Experience in supporting data platforms and ML systems in production settings.

• MLOps experience, including model deployment, monitoring, and retraining workflows.


🏝️ Benefits

• Lead by example, act swiftly, and make informed, data-driven decisions.

• Continuous drive for improvement, always focused on delivering tangible value to customers.

People also viewed

CWILL19 hours ago

DevOps/SRE Engineer, Bilingual Mandarin

US flagCalifornia, +4 more statesFull-timeDevOps & Site Reliability Engineer (SRE)$100k – $130k/year
ApplyView job
a3720 hours ago

Forward Deployed DevOps Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
GT20 hours ago

Site Reliability Engineer, SRE

PL flagPoland, +2 more statesFreelanceDevOps & Site Reliability Engineer (SRE)
ApplyView job
Sigma Software Group21 hours ago

DevOps Engineer

PL flagPoland OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Applaudo21 hours ago

Google Cloud DevOps Engineer – Temporary Contract

CO flagColombia OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Branch21 hours ago

Cloud Operations Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$135k – $150k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers