
Senior DevOps/SRE Engineer
Posted 22 hours ago

Posted 22 hours ago
This is a fully remote position, open to applicants in United Kingdom.
• Design, develop, and manage AWS cloud infrastructure, encompassing IAM, EC2, EKS, networking, databases, serverless, and analytics services.
• Take ownership of Terraform infrastructure-as-code and delivery pipelines throughout various environments.
• Architect systems with a focus on resilience, maintainability, availability, security, observability, capacity, and operability.
• Securely and reliably integrate and configure third-party services.
• Utilize Python and other programming languages to automate DevOps tasks and enhance production capabilities.
• Identify, document, and mitigate infrastructure and platform-level technical risks.
• Design, construct, and sustain efficient, reliable, and secure CI/CD pipelines.
• Develop infrastructure-as-code security and guardrail solutions, including Helm charts, OPA policies, and EKS admission controllers.
• Collaborate with Engineering to enhance build, test, release, and deployment workflows.
• Contribute to strategies for environment and release management.
• Assist Engineers during rollouts and deployments.
• Work alongside Engineering to address application constraints and make platform decisions.
• Analyze application code to support root-cause analysis and informed infrastructure choices.
• Represent the DevOps team in architecture, design, and planning discussions.
• Manage infrastructure and delivery projects from proposal through to delivery and operations.
• Research new tools and practices, proposing enhancements for the platform and operations.
• Assist in defining technical direction and provide support for senior technical decisions as needed.
• Document technical decisions through ADRs/RFCs and contribute to team standards and architectures.
• Create observability using metrics, logs, tracing, and alerting mechanisms.
• Implement SRE practices, lead incident responses and blameless post-incident reviews, and work to eliminate recurring issues.
• Enhance signal quality, reduce alert noise, and address observability gaps.
• Evaluate and strengthen security resilience, IT/network controls, and audit evidence.
• Support corporate IT infrastructure, including Microsoft 365 and ZTNA.
• Assist with SOC 2 and Cyber Essentials Plus compliance evidence, controls, and remediation efforts.
• Participate in on-call rotations and share responsibility for infrastructure and delivery platforms.
• Candidates must be located within the UK.
• No visa sponsorship is available.
• Extensive hands-on experience with infrastructure, including in-depth knowledge of network architecture design, troubleshooting, and reasoning.
• Experience in developing, deploying, and supporting intricate infrastructures.
• Proficient in production-grade Infrastructure as Code, including design, state management, and multi-environment workflows.
• Scripting experience in languages such as Python.
• Comprehensive AWS knowledge in compute, networking, storage, IAM, and managed services in a multi-account environment.
• Strong understanding of access controls, encryption, networking concepts, and protocols.
• Proven experience in designing and operating CI/CD delivery pipelines in a production environment.
• Familiarity with containerization and orchestration technologies like Docker and Kubernetes.
• Experience implementing security and guardrail controls in Infrastructure as Code and Kubernetes settings.
• Knowledge of monitoring and observability tools, including metrics, logging, tracing, and alerting.
• SRE experience or equivalent, focusing on production reliability, incident response, and SLIs/SLOs.
• Recent hands-on experience in developing and supporting modern, multi-tenant systems.
• Experience with AI-assisted tools and workflows in daily engineering tasks.
• Up-to-date with current DevOps practices and collaborative work with developers/engineers.
• Proven track record of managing technical initiatives from design to delivery with minimal oversight.
• Ability to engage as a credible technical peer in senior/lead-level design and architecture discussions.
• Sufficient software development knowledge to read and analyze backend and/or frontend application code.
• Understanding of resilience, security, and operability, including failover, backup/recovery, least-privilege access, secrets management, and observability.
• Excellent written and verbal communication skills.
• Desirable: experience in a small to mid-sized organization with closely integrated DevOps and Engineering teams.
• Desirable: familiarity with ADR and RFC practices.
• Desirable: hands-on experience with Kotlin/Java and/or React development.
• Desirable: knowledge of Microsoft 365 administration and/or ZTNA network solutions.
• Desirable: support for SOC 2 and/or Cyber Essentials Plus certification or renewal.
• Option for fully remote work.
• Flexible working hours.
• Automatic enrollment in the pension scheme with contributions from FXE.
• 25 days of annual leave plus UK Bank Holidays.
• Talent Referral Scheme.
• Regular social events and team gatherings.
Slate Auto
Leidos
LeoLabs
OpenRouter
Get handpicked remote jobs straight to your inbox weekly.