
Cloud Infrastructure Engineer
Posted Jul 25

Posted Jul 25
This is a fully remote position, open to applicants in United Kingdom.
• Designing, constructing, and maintaining AWS infrastructure utilizing Terraform (EC2, RDS, S3, SQS, Lambda, ALB, ElastiCache, Route 53, VPC networking)
• Writing and sustaining Puppet modules to configure and manage fleets of EC2 instances across various auto-scaling groups
• Maintaining and enhancing Python-based automation and tooling that facilitates platform operations
• Operating and refining distributed service discovery and configuration management (etcd)
• Managing and optimizing a multi-tier caching strategy (Varnish, Redis/Valkey, PHP OPcache)
• Running and scaling our observability stack (Prometheus, Grafana, Loki, Fluentd, PagerDuty) and participating in on-call rotations
• Evaluating and implementing distributed storage solutions as the platform progresses
• Enhancing deployment workflows and release processes
• Collaborating with internal teams on API contracts, integration patterns, and operational tooling
• Engaging in incident response, root cause analysis, and platform reliability enhancements
• Extensive experience with AWS services in a production environment — particularly EC2, RDS, S3, SQS, Lambda, ALB, ElastiCache, Route 53, IAM, and VPC networking
• Expertise in authoring and maintaining Terraform modules for production infrastructure
• Proficiency in writing and sustaining Puppet modules (or similar agent-based configuration management) for fleet management
• Strong Python skills — you will be developing and maintaining production daemons, not just scripts
• In-depth knowledge of Linux systems (Ubuntu) — comfortable with Apache/Nginx, PHP-FPM, Varnish, systemd, filesystem mounts, and networking fundamentals
• Understanding of distributed systems concepts: consensus, leader election, distributed locking, eventual consistency, and the associated trade-offs
• Proficient in building and maintaining observability pipelines (Prometheus, Grafana, Loki, or equivalent) in a production setting
• Comfortable working within a GitLab-based CI/CD workflow
• Clear communicator who can document architectural decisions and convey technical trade-offs to both technical and non-technical stakeholders
• Practical experience with distributed storage systems such as Ceph, GlusterFS, JuiceFS, CubeFS, or AWS EFS — especially in the context of migration or evaluation
• Familiar with etcd (or similar distributed key-value stores like Consul or ZooKeeper) including watch APIs, TTL-based locking, and cluster operations
• Experience with Varnish and VCL, particularly dynamic backend routing or multi-tenant configurations
• Working knowledge of PHP — not for application development, but to understand and maintain integration scripts that connect infrastructure and application layers
• Background in multi-tenant SaaS platform design — particularly database-per-tenant models on shared infrastructure
• Familiarity with Moodle LMS or educational technology platforms
• Experience with secrets management solutions (AWS Secrets Manager, HashiCorp Vault, Parameter Store) and automated credential rotation
• Experience in designing zero-downtime deployment strategies for VM-based (non-containerized) environments
• Equal employment opportunity
• Affirmative action employer
Tailscale
Adzuna
Webflow
Coinbase
Get handpicked remote jobs straight to your inbox weekly.