
DevOps/SRE Engineer
Posted Jul 27

Posted Jul 27
This is a fully remote position, open to applicants in United Kingdom.
• Design, develop, and maintain cloud infrastructure utilizing AWS technologies such as EKS, EC2, RDS, CloudWatch, and IAM.
• Oversee and optimize the performance and reliability of Kubernetes (EKS) clusters.
• Create and refine CI/CD pipelines alongside deployment automation processes.
• Establish observability frameworks including centralized logging, metrics dashboards, alerting, and distributed tracing.
• Define and monitor service level indicators (SLIs), service level objectives (SLOs), and error budgets across platform services.
• Lead the processes for incident management and response; facilitate post-incident analysis.
• Collaborate with engineering teams to minimize failure points and support self-healing systems.
• Contribute to the security and compliance posture according to FCA and PCI DSS standards.
• Drive initiatives focused on infrastructure cost optimization and operational efficiency.
• 5–7+ years of experience in DevOps, Site Reliability Engineering (SRE), or Platform Engineering roles.
• Strong practical experience with AWS and Kubernetes, preferably EKS.
• Demonstrated ability in building and maintaining CI/CD pipelines, with a preference for GitLab CI.
• Experience with Infrastructure as Code, such as Terraform or similar tools.
• Comprehensive understanding of distributed systems, high availability, and fault tolerance principles.
• Experience with monitoring, alerting, and logging systems in live production environments.
• Ability to configure and manage databases, closely aligned with DBA responsibilities.
• Familiarity with networking fundamentals and best practices in cloud security.
• Strong sense of ownership, capable of working independently and taking full responsibility for projects.
• Nice to Have: Experience in fintech, payment systems, or other regulated sectors; Knowledge of PCI DSS compliance requirements; Practical experience with OpenSearch/ELK, Prometheus/Grafana, or OpenTelemetry; Interest or experience in AIOps or machine learning-based monitoring and anomaly detection; Exposure to event-driven architectures, including Kafka or AMQP.
• Competitive compensation that aligns with market standards, discussed openly during the first interview.
• Full ownership of platform reliability for a live fintech product, not just within a sandbox environment.
• Opportunity to establish SRE practices and observability tools from the ground up.
• Exposure to a modern cloud-native and observability technology stack.
• Clear growth trajectory towards Senior SRE, Platform Lead, or Cloud Architect roles.
• Fully remote working environment with flexible hours, alongside a highly skilled international team.
CWILL
a37
GT
Sigma Software Group
Get handpicked remote jobs straight to your inbox weekly.