Senior DevOps Engineer

Posted Sep 17

This is a fully remote position, open to applicants in Philippines.

📋 Description

• Take ownership of and continuously enhance cloud infrastructure and platform capabilities that support applications, data systems, and engineering teams.

• Identify areas for improvement in reliability, scalability, security, performance, cost efficiency, developer experience, and operational effectiveness.

• Design, construct, and maintain secure and scalable cloud infrastructure utilizing Infrastructure as Code and cloud-native methodologies.

• Develop automation, self-service options, reusable infrastructure patterns, and platform tools to minimize repetitive operational tasks.

• Enhance Kubernetes and container-based infrastructure to ensure standardized and dependable deployment, scaling, recovery, and operations.

• Own and refine CI/CD and GitOps workflows.

• Collaborate with Software Engineers, Product Engineers, Analytics Engineers, and other Technology teams to create platform solutions.

• Boost observability through monitoring, logging, tracing, alerting, dashboards, and actionable service health indicators.

• Lead technical investigations during complex production incidents, driving systemic enhancements.

• Strengthen resilience, disaster recovery, backup strategies, security controls, access management, and infrastructure risk assessment.

• Assess infrastructure and architecture for cost efficiency and work to eliminate unnecessary expenses.

• Utilize AI-assisted engineering in infrastructure development, troubleshooting, root cause analysis, automation, documentation, and operational workflows.

• Contribute to architectural decisions and the long-term platform strategy.

• Define, document, and advocate for engineering standards, operational practices, reusable patterns, and platform capabilities.

• Mentor engineers and elevate standards for infrastructure ownership, automation, reliability, and operational excellence.


⛳️ Requirements

• Proven ability to independently identify infrastructure and operational issues, ascertain root causes, design solutions, and implement improvements yielding measurable outcomes.

• Established track record managing business-critical production infrastructure and complex platform initiatives with minimal supervision.

• Strong judgment in balancing reliability, security, engineering velocity, cost, complexity, and business requirements.

• Capability to operate effectively during production incidents.

• Proficiency in utilizing AI-assisted engineering tools for infrastructure development, troubleshooting, automation, documentation, incident analysis, and enhancing productivity.

• Ability to critically assess AI-generated infrastructure and operational modifications prior to production deployment.

• Extensive experience in designing, operating, troubleshooting, and enhancing highly available production systems within cloud environments.

• Solid expertise with Kubernetes, containerization, and contemporary cloud-native infrastructure.

• Strong experience with Infrastructure as Code using Terraform, OpenTofu, or similar technologies.

• Comprehensive understanding of CI/CD, GitOps, automated deployment practices, and modern software delivery workflows.

• Experience with production observability, including monitoring, logging, tracing, alerting, and incident management.

• Strong grasp of reliability engineering, failure modes, resilience, capacity, recovery, and operational risk.

• Familiarity with secure infrastructure patterns, access controls, secrets management, system hardening, and cloud security.

• Proficient in Linux and systems troubleshooting.

• Significant experience managing production workloads in AWS or a comparable cloud platform.

• Experience with GitHub Actions, ArgoCD, or similar platforms.

• Familiarity with Datadog or comparable observability platforms.

• Strong scripting or programming skills in Python, Bash, or another suitable language.

• Working knowledge of networking, DNS, CDN/WAF technologies, load balancing, TLS, and application delivery architecture.

• Experience supporting MySQL, Redis, Redshift, or similar data stores and stateful systems.

• Excellent communication skills with both technical and non-technical stakeholders.

• Ability to collaborate effectively with Software Engineering, Analytics Engineering, Product, Security, and other teams.

• Self-driven, highly autonomous, and comfortable in a remote-first setting.

• Ability to constructively challenge existing approaches and influence technical direction.

• Capacity to mentor engineers and elevate engineering standards.

• Authorized to work in the Philippines at hire and throughout employment.

• Pre-employment screening required.


🏝️ Benefits

• Mentorship programs.

• Leadership academies.

• Opportunities to shape company culture and DEI initiatives.

• Flex Fridays: adjust the 40-hour workweek to enjoy a full or half day off on Fridays.

• Remote-first culture.

• 14 days of annual paid time off.

• Philippine government-declared holidays.

• 5 additional paid time-off days after one year.

• Philippine statutory benefits: SSS, PhilHealth, and HDMF.

• Corporate HMO Plan for full-time PH employees after two months of tenure.

• Monthly rice subsidy.

• Access to the Headspace app.

• Speaker Series bonus.

People also viewed

Horizon3.ai1 day ago

Staff Site Reliability Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$199.8k – $270k/year
ApplyView job
CLOUD MANTA GmbH1 day ago

Senior DevOps Engineer, Containers & Private Cloud

DE flagGermany OnlyFull-timeDevOps & Site Reliability Engineer (SRE)€70k – €80k/year
ApplyView job
Stefanini LATAM1 day ago

Senior DevOps

AR flagArgentina OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Akamai Technologies1 day ago

Principal Site Reliability Engineer – Lead

PL flagPoland OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
PingWind Inc. (SDVOSB)1 day ago

DevSecOps Engineer

US flagAlabama, +1 more stateFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Ad Hoc LLC1 day ago

Staff DevOps Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$130k – $150k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers