
Infrastructure & Platform Operations Engineer
Posted Aug 17

Posted Aug 17
This is a fully remote position, open to applicants in Philippines.
β’ Assist in the support, maintenance, and enhancement of infrastructure and production platforms utilized to deliver Nephos services to clients.
β’ Provide business-as-usual (BAU) production support, ensuring service stability and knowledge transfer across both cloud and on-premise infrastructure.
β’ Oversee infrastructure and applications, investigate alerts and logs, identify root causes, and minimize unnecessary alert noise.
β’ Administer and troubleshoot customer environments, including Linux systems, networking, PostgreSQL, MongoDB, Docker, Kubernetes, and AWS container services.
β’ Utilize Python, PowerShell, and other scripting and automation tools to increase repeatability and reduce manual tasks.
β’ Manage incidents, problems, service requests, and changes, which includes troubleshooting, service restoration, root-cause analysis, validation, and rollback planning.
β’ Support enterprise platform installations, configurations, upgrades, testing, validation, and recovery efforts.
β’ Create and maintain knowledge articles, runbooks, and work instructions while participating in cross-training initiatives.
β’ Collaborate with customers and internal technical teams during incidents, changes, upgrades, and investigations.
β’ Work alongside DevOps, Data Services, Product, Delivery, and other cross-functional teams to enhance collaboration.
β’ Contribute to improvements in platform resilience, capacity, monitoring, and operational design.
β’ Assist with backup, restore, disaster recovery activities, and practical recovery exercises.
β’ Implement security principles such as least privilege, secure credential handling, secrets management, certificates, auditability, patching, and vulnerability awareness.
β’ Typically requires 5+ years of experience in infrastructure engineering, platform operations, systems engineering, or production support; demonstrable capability is prioritized over a specific number of years.
β’ Extensive BAU and production support experience, including handling live incidents, service restoration, monitoring, failed changes or deployments, root-cause analysis, and controlled remediation.
β’ Strong hands-on experience with AWS in production environments; AWS is the primary cloud requirement.
β’ Solid Linux administration and troubleshooting skills, especially with Ubuntu and Red Hat, including command-line proficiency.
β’ Profound networking and connectivity troubleshooting experience, encompassing TCP/IP, DNS, routing, firewalls, VPNs, TLS, and cloud networking.
β’ Strong production-level MongoDB administration and troubleshooting experience.
β’ Significant PostgreSQL operational experience, including administration, backup and recovery, monitoring, and troubleshooting.
β’ Hands-on experience with Docker in production settings and a solid understanding of the container lifecycle, networking, logging, health, and troubleshooting.
β’ Strong hands-on experience with Kubernetes, including deployment health, pods, services, configuration, logs, and failure diagnosis.
β’ Experience in supporting enterprise applications and platforms through installation, configuration, upgrades, health monitoring, and technical troubleshooting.
β’ Strong scripting and automation skills, particularly in Python and/or PowerShell; familiarity with Bash or other transferable scripting languages is relevant.
β’ Strong fundamentals in monitoring and observability, including infrastructure metrics, log investigation, and alert analysis.
β’ Working knowledge of Git and GitHub.
β’ Practical understanding of backup, restore, recovery, and service resilience.
β’ Relevant IT service management experience encompassing Incident, Problem, Change, Service Request, and major-incident processes.
β’ Experience with service management platforms such as Jira/JSM, ServiceNow, Remedy, Freshservice, or equivalent tools.
β’ Good security awareness, covering least privilege, IAM concepts, secrets management, certificates, MFA, privileged access, patching, and vulnerability awareness.
β’ Strong customer-centric attitude with effective communication skills for both technical and non-technical stakeholders.
β’ Excellent written and verbal communication skills, including proficiency in technical documentation and operational procedures.
β’ Ability to learn complex unfamiliar technologies and become independently effective after onboarding and knowledge transfer.
β’ Capability to work autonomously, manage multiple priorities, and recognize when escalation or broader technical input is required.
β’ Willingness to participate in occasional out-of-hours upgrades, changes, and major incidents.
β’ Nice-to-have: Familiarity with Azure, Google Cloud Platform, infrastructure as code, CI/CD, BigID, RabbitMQ, Redis, LogicMonitor, privileged-access and secrets management technologies, REST APIs, TLS certificates, disaster recovery planning, data governance, relevant certifications, microservices, serverless, or distributed architectures.
β’ Government-mandated benefits in addition to supplemental HMO coverage.
β’ A collaborative and professional work environment.
β’ Opportunities for career growth within an expanding technology organization.
β’ Hybrid/Remote work arrangements where applicable.
β’ Competitive compensation package based on experience.
Quvia
Anyone AI
Manulife
Accenture Federal Services
Get handpicked remote jobs straight to your inbox weekly.