
Cloud Infrastructure Engineer
Posted 22 hours ago

Posted 22 hours ago
This is a fully remote position, open to applicants in Colombia.
• Design, develop, and manage AWS infrastructure utilizing Terraform.
• Create and sustain Puppet modules for numerous EC2 instance fleets spanning various auto-scaling groups.
• Maintain and enhance Python-driven automation and tools that support platform operations.
• Operate and refine distributed service discovery and configuration management with etcd.
• Oversee and optimize caching layers including Varnish, Redis/Valkey, and PHP OPcache.
• Deploy and scale observability systems such as Prometheus, Grafana, Loki, Fluentd, and PagerDuty.
• Engage in on-call rotations.
• Assess and deploy distributed storage solutions.
• Enhance deployment workflows and release processes.
• Collaborate with internal teams on API agreements, integration methods, and operational tools.
• Take part in incident response, root cause analysis, and improvements to platform reliability.
• Assist in building, scaling, and evolving Open LMS's multi-tenant SaaS hosting platform on AWS for Moodle LMS instances.
• Extensive experience with AWS services in a production setting, especially EC2, RDS, S3, SQS, Lambda, ALB, ElastiCache, Route 53, IAM, and VPC networking.
• Expertise in creating and managing Terraform modules for production infrastructure.
• Proficient in writing and maintaining Puppet modules or equivalent agent-based configuration management for fleet oversight.
• Strong Python programming skills, including experience with production daemons.
• Comprehensive knowledge of Linux systems, particularly Ubuntu.
• Familiarity with Apache/Nginx, PHP-FPM, Varnish, systemd, filesystem mounts, and networking basics.
• Understanding of distributed systems principles, including consensus, leader election, distributed locking, eventual consistency, and their trade-offs.
• Proficient in production observability pipelines using Prometheus, Grafana, Loki, or similar tools.
• Comfort with a GitLab-based CI/CD workflow.
• Strong communication skills and the ability to document architectural decisions and articulate technical trade-offs to both technical and non-technical audiences.
• Hands-on experience with distributed storage systems like Ceph, GlusterFS, JuiceFS, CubeFS, or AWS EFS is preferred.
• Familiarity with etcd or similar distributed key-value stores, including watch APIs, TTL-based locking, and cluster operations is preferred.
• Experience with Varnish and VCL is preferred.
• Working knowledge of PHP is preferred.
• Background in the design of multi-tenant SaaS platforms is preferred.
• Familiarity with Moodle LMS or educational technology platforms is preferred.
• Experience with AWS Secrets Manager, HashiCorp Vault, Parameter Store, and automated credential rotation is preferred.
• Experience in designing zero-downtime deployment strategies for VM-based, non-containerized environments is preferred.
• Equal employment opportunity/affirmative action employer.
• Full-time employment.
ZoomInfo
Skaylink
Kyndryl
Milliman
Get handpicked remote jobs straight to your inbox weekly.