
Senior DevOps Engineer
Posted Sep 17

Posted Sep 17
This is a fully remote position, open to applicants in Philippines.
• Take ownership of and continuously enhance cloud infrastructure and platform capabilities that support applications, data systems, and engineering teams.
• Identify areas for improvement in reliability, scalability, security, performance, cost efficiency, developer experience, and operational effectiveness.
• Design, construct, and maintain secure and scalable cloud infrastructure utilizing Infrastructure as Code and cloud-native methodologies.
• Develop automation, self-service options, reusable infrastructure patterns, and platform tools to minimize repetitive operational tasks.
• Enhance Kubernetes and container-based infrastructure to ensure standardized and dependable deployment, scaling, recovery, and operations.
• Own and refine CI/CD and GitOps workflows.
• Collaborate with Software Engineers, Product Engineers, Analytics Engineers, and other Technology teams to create platform solutions.
• Boost observability through monitoring, logging, tracing, alerting, dashboards, and actionable service health indicators.
• Lead technical investigations during complex production incidents, driving systemic enhancements.
• Strengthen resilience, disaster recovery, backup strategies, security controls, access management, and infrastructure risk assessment.
• Assess infrastructure and architecture for cost efficiency and work to eliminate unnecessary expenses.
• Utilize AI-assisted engineering in infrastructure development, troubleshooting, root cause analysis, automation, documentation, and operational workflows.
• Contribute to architectural decisions and the long-term platform strategy.
• Define, document, and advocate for engineering standards, operational practices, reusable patterns, and platform capabilities.
• Mentor engineers and elevate standards for infrastructure ownership, automation, reliability, and operational excellence.
• Proven ability to independently identify infrastructure and operational issues, ascertain root causes, design solutions, and implement improvements yielding measurable outcomes.
• Established track record managing business-critical production infrastructure and complex platform initiatives with minimal supervision.
• Strong judgment in balancing reliability, security, engineering velocity, cost, complexity, and business requirements.
• Capability to operate effectively during production incidents.
• Proficiency in utilizing AI-assisted engineering tools for infrastructure development, troubleshooting, automation, documentation, incident analysis, and enhancing productivity.
• Ability to critically assess AI-generated infrastructure and operational modifications prior to production deployment.
• Extensive experience in designing, operating, troubleshooting, and enhancing highly available production systems within cloud environments.
• Solid expertise with Kubernetes, containerization, and contemporary cloud-native infrastructure.
• Strong experience with Infrastructure as Code using Terraform, OpenTofu, or similar technologies.
• Comprehensive understanding of CI/CD, GitOps, automated deployment practices, and modern software delivery workflows.
• Experience with production observability, including monitoring, logging, tracing, alerting, and incident management.
• Strong grasp of reliability engineering, failure modes, resilience, capacity, recovery, and operational risk.
• Familiarity with secure infrastructure patterns, access controls, secrets management, system hardening, and cloud security.
• Proficient in Linux and systems troubleshooting.
• Significant experience managing production workloads in AWS or a comparable cloud platform.
• Experience with GitHub Actions, ArgoCD, or similar platforms.
• Familiarity with Datadog or comparable observability platforms.
• Strong scripting or programming skills in Python, Bash, or another suitable language.
• Working knowledge of networking, DNS, CDN/WAF technologies, load balancing, TLS, and application delivery architecture.
• Experience supporting MySQL, Redis, Redshift, or similar data stores and stateful systems.
• Excellent communication skills with both technical and non-technical stakeholders.
• Ability to collaborate effectively with Software Engineering, Analytics Engineering, Product, Security, and other teams.
• Self-driven, highly autonomous, and comfortable in a remote-first setting.
• Ability to constructively challenge existing approaches and influence technical direction.
• Capacity to mentor engineers and elevate engineering standards.
• Authorized to work in the Philippines at hire and throughout employment.
• Pre-employment screening required.
• Mentorship programs.
• Leadership academies.
• Opportunities to shape company culture and DEI initiatives.
• Flex Fridays: adjust the 40-hour workweek to enjoy a full or half day off on Fridays.
• Remote-first culture.
• 14 days of annual paid time off.
• Philippine government-declared holidays.
• 5 additional paid time-off days after one year.
• Philippine statutory benefits: SSS, PhilHealth, and HDMF.
• Corporate HMO Plan for full-time PH employees after two months of tenure.
• Monthly rice subsidy.
• Access to the Headspace app.
• Speaker Series bonus.
Horizon3.ai
CLOUD MANTA GmbH
Stefanini LATAM
Akamai Technologies
Get handpicked remote jobs straight to your inbox weekly.