Site Reliability Engineer

Posted Aug 18

This is a fully remote position, open to applicants in Spain.

📋 Description

• Establish and execute SLIs, SLOs, and error budgets for essential CloudBlue services

• Shape system architecture with an emphasis on reliability, scalability, and operability

• Minimize operational toil through automation and process enhancements

• Construct and manage the observability stack encompassing metrics, logs, and traces

• Create alerting strategies and dashboards for monitoring platform and business health

• Design and sustain high-availability architectures, alongside redundancy, failover, and disaster recovery plans

• Perform capacity planning, load testing, and optimization of performance

• Oversee production incident coordination, communication, and service recovery

• Conduct blameless postmortems and advocate for improvements to decrease incidents, MTTR, and customer impact

• Enhance the reliability of Kubernetes-based platforms through health checks, autoscaling, rollout safety, and resilience testing

• Collaborate with Engineering and DevOps teams on deployment safety, rollback strategies, and platform reliability

• Keep runbooks and operational documentation up to date

• Foster SRE best practices across engineering teams

• Assist with additional tasks or projects as assigned


⛳️ Requirements

• 3+ years of experience in an SRE, DevOps Engineer, or Production Engineer role

• Strong sense of ownership regarding production systems

• Experience in managing highly available, enterprise-grade, multi-tenant SaaS platforms

• Proficient in Datadog, Grafana, and Elasticsearch/Kibana

• Strong foundation in Linux, networking, and distributed systems principles

• Familiarity with Docker and Kubernetes

• Strong scripting and automation capabilities in Python and/or Bash

• Experience in participating in on-call rotations and responding to production incidents

• Excellent written and verbal communication skills in English

• Experience in cloud environments, preferably Azure; experience with AWS and/or GCP is also valued

• Beneficial experience with hybrid or on-premises integrations

• Advantageous experience in defining SLIs/SLOs and managing error budgets, hyperscale/service-provider-grade platforms, chaos engineering, and resilience testing


🏝️ Benefits

• A competitive salary that recognizes your unique skills and contributions

• Opportunities for career advancement and professional development

• Flexible work arrangements to promote work/life balance

• 24/7 award-winning customer support

• Commitment to diversity and inclusion

• Accommodation may be provided throughout all stages of the hiring process

People also viewed

Entarian1 day ago

DevOps Engineer

US flagVirginia OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$150k – $200k/year
ApplyView job
BeyondTrust1 day ago

Senior Manager, Site Reliability Engineer – FedRAMP, AWS GovCloud

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
BeyondTrust1 day ago

Site Reliability Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Scribe1 day ago

Senior DevOps Engineer

US flagCalifornia OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$150k – $240k/year
ApplyView job
Hypertegrity AG1 day ago

Senior DevOps Engineer – Smart City Open Source

DE flagGermany OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
CACI International Inc1 day ago

Senior DevSecOps Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$98.5k – $206.8k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers