Remotery

Site Reliability Engineer

Posted 2 days ago

This is a fully remote position, open to applicants in Spain.

📋 Description

• Establish and execute SLIs, SLOs, and error budgets for essential CloudBlue services

• Shape system architecture with an emphasis on reliability, scalability, and operability

• Minimize operational toil through automation and process enhancements

• Construct and manage the observability stack encompassing metrics, logs, and traces

• Create alerting strategies and dashboards for monitoring platform and business health

• Design and sustain high-availability architectures, alongside redundancy, failover, and disaster recovery plans

• Perform capacity planning, load testing, and optimization of performance

• Oversee production incident coordination, communication, and service recovery

• Conduct blameless postmortems and advocate for improvements to decrease incidents, MTTR, and customer impact

• Enhance the reliability of Kubernetes-based platforms through health checks, autoscaling, rollout safety, and resilience testing

• Collaborate with Engineering and DevOps teams on deployment safety, rollback strategies, and platform reliability

• Keep runbooks and operational documentation up to date

• Foster SRE best practices across engineering teams

• Assist with additional tasks or projects as assigned


⛳️ Requirements

• 3+ years of experience in an SRE, DevOps Engineer, or Production Engineer role

• Strong sense of ownership regarding production systems

• Experience in managing highly available, enterprise-grade, multi-tenant SaaS platforms

• Proficient in Datadog, Grafana, and Elasticsearch/Kibana

• Strong foundation in Linux, networking, and distributed systems principles

• Familiarity with Docker and Kubernetes

• Strong scripting and automation capabilities in Python and/or Bash

• Experience in participating in on-call rotations and responding to production incidents

• Excellent written and verbal communication skills in English

• Experience in cloud environments, preferably Azure; experience with AWS and/or GCP is also valued

• Beneficial experience with hybrid or on-premises integrations

• Advantageous experience in defining SLIs/SLOs and managing error budgets, hyperscale/service-provider-grade platforms, chaos engineering, and resilience testing


🏝️ Benefits

• A competitive salary that recognizes your unique skills and contributions

• Opportunities for career advancement and professional development

• Flexible work arrangements to promote work/life balance

• 24/7 award-winning customer support

• Commitment to diversity and inclusion

• Accommodation may be provided throughout all stages of the hiring process

People also viewed

CWILL15 hours ago

DevOps/SRE Engineer, Bilingual Mandarin

US flagCalifornia, +4 more statesFull-timeDevOps & Site Reliability Engineer (SRE)$100k – $130k/year
ApplyView job
a3716 hours ago

Forward Deployed DevOps Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
GT17 hours ago

Site Reliability Engineer, SRE

PL flagPoland, +2 more statesFreelanceDevOps & Site Reliability Engineer (SRE)
ApplyView job
Sigma Software Group17 hours ago

DevOps Engineer

PL flagPoland OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Applaudo17 hours ago

Google Cloud DevOps Engineer – Temporary Contract

CO flagColombia OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Branch17 hours ago

Cloud Operations Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$135k – $150k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers