
Principal Platform Engineer
Posted Jul 28

Posted Jul 28
This is a fully remote position, open to applicants in United States.
• Oversee the design, development, deployment, and operational management of automated, resilient, high-availability, self-healing, and secure platforms equipped with native-AI capabilities for IT requirements, catering to both internal and customer business functionalities.
• Lead, build, and manage the Platform Engineering team and function, which includes hiring, mentoring, performance management, and ownership of the technical roadmap.
• Strategize, construct, and maintain an OpenTelemetry Observability platform utilizing technologies such as Grafana, Mimir, Loki, Tempo, and Alertmanager on Kubernetes/RKE2, implemented with Helm and ArgoCD.
• Develop an automated federated Observability Edge Stack comprising Prometheus and OTel collector nodes deployed per site, along with Zabbix auto-discovery configurations and a Prometheus scrape profile library for over 10 device classes (Cisco, Juniper, Dell, NetApp, etc.).
• Design, develop, and manage engineering lifecycle platforms to facilitate high-velocity secure SDLC using GitLab and related technologies.
• Construct and operate infrastructure as code (IaC) and CI/CD platforms, including GitLab CI/CD, Terraform, Ansible AWX, Helm, and ArgoCD for automated provisioning and application deployment.
• Own, enhance, and manage critical IT platform technologies, such as Boomi for integrations and AWS for Cloud environments, including their hosted infrastructure.
• Establish and uphold platform security measures, including secrets management via CyberArk/Conjur, RBAC, mTLS, compliance boundary design, and a zero inbound telemetry architecture.
• Develop and integrate ITSM capabilities across various platforms, such as automated incident creation, CI enrichment, and CMDB correlation.
• Define and implement extensibility patterns, including AIOps, such as anomaly detection hooks, event correlation pipeline design, and integration with future ML/AI tools.
• Collaborate with other IT and business teams for application development, requirements capture, delivery validation, and integration purposes.
• Represent platform engineering during cross-functional architecture reviews and executive-level program updates.
• Perform additional management and technical responsibilities as needed and assigned to ensure team and operational resilience, including team building and on-call rotation.
• Travel may be necessary for team or project events.
• A minimum of 8 years of relevant technical experience, including at least 2 years in a management (or Principal-level) role leading an engineering team.
• DevOps/Platform Engineering experience of 8+ years, with end-to-end ownership of developer/infrastructure platforms, including Kubernetes, Helm, ArgoCD, service mesh, and containerized workloads.
• GitOps/CI-CD experience of 5+ years, with proficiency in GitLab CI/CD, pipeline authoring, and infrastructure-as-code delivery.
• More than 8 years of advanced automation framework experience with Python, Terraform, Ansible, etc.
• 8+ years of experience in infrastructure, particularly in Linux systems administration, VM lifecycle management (VMware vCenter/VCF), and NetApp storage and compute provisioning.
• Working knowledge of Networking for 3+ years, including TCP/IP, BGP/OSPF, and SNMP protocols.
• Strong understanding (or 1+ years of experience) with AI tooling, including MCP, Agentic workflows, and SRE workflows such as AIOps for anomaly detection, event correlation, and alert noise reduction in the Prometheus and Grafana stack.
• At least 4 years of experience with Secrets & Security, including CyberArk, Conjur, Vault, or similar; experience in RBAC design and compliance boundary architecture.
• Minimum of 4 years in Engineering Management, including hiring, team building, performance management, and roadmap ownership for teams of 5 or more engineers.
• Other relevant training and experience may be considered in lieu of job requirements at the discretion of the manager.
• Medical, Telehealth, Dental, and Vision coverage
• 401(k) plan
• Health Savings Accounts (HSA) and Flexible Spending Accounts (FSA)
• Life and AD&D insurance
• Short Term and Long-Term disability coverage
• Flexible Paid Time Off (PTO)
• Leave of Absence options
• Employee Assistance Program
• Wellness Program
• Rewards and Recognition Program
Quantiphi
Encompass Corporation
Shippit
Group O
Get handpicked remote jobs straight to your inbox weekly.