
Cloud Engineer, Azure Platform Engineer
Posted Jul 30

Posted Jul 30
This is a fully remote position, open to applicants in Canada.
• Serve as the subject-matter expert (SME) for Kubernetes deployments, addressing troubleshooting and production challenges across all environments.
• Design and implement Azure Kubernetes Service (AKS) clusters featuring private cluster configurations, managed identities, and role-based access control (RBAC).
• Take ownership of the health, scaling, and lifecycle management of production AKS clusters, including upgrades, node pool management, autoscaling, and capacity planning.
• Configure AKS networking components, such as Azure CNI, internal load balancers, and ingress controllers (NGINX, Traefik).
• Design and sustain integrations between AKS and Azure Container Registry (ACR), Key Vault through CSI driver, Azure Monitor for containers, Azure SQL, Kafka/Event Hubs, Azure Storage, along with other client-facing dependencies (DNS resolution, firewall rules, private endpoints, and service connectivity).
• Manage the deployment of containerized applications using Docker and Helm; uphold standards for reusable charts and templating, namespaces, resource quotas, and Azure Policy for AKS.
• Strengthen AKS environments via policy enforcement, network policies, and image scanning.
• Oversee container and cluster vulnerability management, including scanning, triage, prioritization, and facilitating remediation with engineering teams.
• Contribute to the development of Terraform-based infrastructure as code for the provisioning and management of Azure resources.
• Support Azure DevOps (or similar) CI/CD pipelines, including GitOps workflows (Flux/ArgoCD), to minimize deployment risks and enhance release velocity.
• Design, implement, and oversee observability across the Grafana stack (Prometheus, Loki, Tempo) and Azure-native tools (Azure Monitor, Log Analytics Workspace, Application Insights) for critical systems; establish SLIs/SLOs and adjust alerting to mitigate noise.
• Collaborate with performance engineering and application teams to pinpoint, diagnose, and resolve performance bottlenecks across the AKS platform and its dependent services (database, messaging, network) to appropriately size node pools and workloads based on observed performance and utilization patterns.
• Contribute to Disaster Recovery and Business Continuity Planning (DR/BCP) protocols for AKS-hosted workloads, including participation in cross-team failover drills and RTO/RPO validation.
• Provide escalation support for production Kubernetes and infrastructure incidents; engage in on-call rotation, lead root-cause analyses, and drive preventative follow-up actions.
• Document runbooks and post-incident reviews, along with operational knowledge; maintain an updated knowledge base to minimize tribal knowledge.
• A minimum of 5 years of practical Kubernetes experience, with at least 2 years managing AKS in production environments.
• In-depth knowledge of Kubernetes internals including scheduling, networking, storage, and RBAC.
• Proficient in Azure CNI networking and AKS private cluster configurations.
• Hands-on experience integrating AKS with Azure PaaS services (ACR, AKV, Azure SQL, Kafka, and Managed Identities) and troubleshooting network-layer dependencies (DNS, firewall, private endpoints).
• Experience with Helm and GitOps workflows (Flux/ArgoCD).
• Familiarity with Terraform for infrastructure as code.
• Working knowledge of Azure DevOps or comparable CI/CD tools.
• Practical experience in implementing and managing observability platforms, including the Grafana stack (Prometheus, Loki, Tempo) and Azure-native monitoring (Azure Monitor, Log Analytics, Application Insights); experience in defining SLIs/SLOs for production systems.
• Proven experience with container and cluster vulnerability management and remediation processes.
• Scripting abilities in Bash and Python.
• Required certifications: Certified Kubernetes Administrator (CKA) or CKAD certification, and Microsoft Certified: Azure Administrator (AZ-104).
• Exceptional written and verbal communication skills; capable of clearly conveying technical issues to both technical and non-technical audiences.
• Remote Work Environment
• Flexible Time Away From Work Policy including PTO, Personal and Sick Days
• Competitive Salary and Health/Medical Benefits
• RRSP/TFSA/401K Employee Contribution
• Life and Disability Insurance
• Employee Assistance Program
• FHIR Study Program and Skillsoft Learning
• Super HAPI Fun Club
LouisianaNOW.Jobs
Spassu
Kubermatic
Get handpicked remote jobs straight to your inbox weekly.