
Site Reliability Engineer
Posted 2 days ago

Posted 2 days ago
This is a fully remote position, open to applicants in Spain.
• Establish and execute SLIs, SLOs, and error budgets for essential CloudBlue services
• Shape system architecture with an emphasis on reliability, scalability, and operability
• Minimize operational toil through automation and process enhancements
• Construct and manage the observability stack encompassing metrics, logs, and traces
• Create alerting strategies and dashboards for monitoring platform and business health
• Design and sustain high-availability architectures, alongside redundancy, failover, and disaster recovery plans
• Perform capacity planning, load testing, and optimization of performance
• Oversee production incident coordination, communication, and service recovery
• Conduct blameless postmortems and advocate for improvements to decrease incidents, MTTR, and customer impact
• Enhance the reliability of Kubernetes-based platforms through health checks, autoscaling, rollout safety, and resilience testing
• Collaborate with Engineering and DevOps teams on deployment safety, rollback strategies, and platform reliability
• Keep runbooks and operational documentation up to date
• Foster SRE best practices across engineering teams
• Assist with additional tasks or projects as assigned
• 3+ years of experience in an SRE, DevOps Engineer, or Production Engineer role
• Strong sense of ownership regarding production systems
• Experience in managing highly available, enterprise-grade, multi-tenant SaaS platforms
• Proficient in Datadog, Grafana, and Elasticsearch/Kibana
• Strong foundation in Linux, networking, and distributed systems principles
• Familiarity with Docker and Kubernetes
• Strong scripting and automation capabilities in Python and/or Bash
• Experience in participating in on-call rotations and responding to production incidents
• Excellent written and verbal communication skills in English
• Experience in cloud environments, preferably Azure; experience with AWS and/or GCP is also valued
• Beneficial experience with hybrid or on-premises integrations
• Advantageous experience in defining SLIs/SLOs and managing error budgets, hyperscale/service-provider-grade platforms, chaos engineering, and resilience testing
• A competitive salary that recognizes your unique skills and contributions
• Opportunities for career advancement and professional development
• Flexible work arrangements to promote work/life balance
• 24/7 award-winning customer support
• Commitment to diversity and inclusion
• Accommodation may be provided throughout all stages of the hiring process
CWILL
a37
GT
Sigma Software Group
Get handpicked remote jobs straight to your inbox weekly.