
Site Reliability Engineer
Posted Jun 23

Posted Jun 23
This is a fully remote position, open to applicants in United Kingdom.
• Design and sustain tools aimed at enhancing engineering productivity, including but not limited to the automation of deployment infrastructure and database upgrades.
• Identify processes that can benefit from automation, prioritizing internal Developer Experience.
• Collaborate with engineers to facilitate more efficient workflows.
• Develop, maintain, and test our system's disaster recovery procedures, incorporating tools to automate these processes.
• Manage production incidents, create blameless postmortems, and enhance operational playbooks and runbooks.
• Conduct performance investigations and lead optimization efforts across applications, data stores, and AWS.
• Oversee cost optimization initiatives, including right-sizing, autoscaling policies, workload scheduling, storage tiering, and identifying inefficiencies in ECS/Fargate, RDS/Aurora, SQS, and observability.
• Champion the GitOps methodology.
• A background in software development, with hands-on experience in delivering and managing production services.
• Proficiency in navigating the AWS ecosystem.
• An interest in AI and its transformative impact on software development.
• A collaborative spirit, with a focus on security and attention to detail.
• Familiarity with platform and operations concepts, including networking and Linux administration.
• Experience with microservices and distributed systems at scale.
• Proficient with monitoring tools; our stack includes Opentelemetry, Honeycomb, Grafana, Pingdom, and Incident.io.
• Unlimited vacation days.
• Four months of paid family leave, regardless of gender.
• Annual learning budget.
• Ongoing internal and external training opportunities.
• Options for remote work.
The Codest
IRIUM
Sólides
Resilinc
Get handpicked remote jobs straight to your inbox weekly.