
Azure Platform Engineer
Posted 18 hours ago

Posted 18 hours ago
This is a fully remote position, open to applicants in Canada.
• Serve as the subject-matter expert (SME) for Kubernetes deployments, troubleshooting, and production issues across all environments.
• Design and implement Azure Kubernetes Service (AKS) clusters featuring private cluster configurations, managed identities, and RBAC.
• Take ownership of the health, scaling, and lifecycle management of production AKS clusters, including upgrades, node pool management, autoscaling, and capacity planning.
• Configure AKS networking elements such as Azure CNI, internal load balancers, and ingress controllers (NGINX, Traefik).
• Design and oversee integrations between AKS and Azure Container Registry, Key Vault via CSI driver, Azure Monitor for containers, Azure SQL, Kafka/Event Hubs, Azure Storage, and client-facing dependencies.
• Manage deployments of containerized applications using Docker and Helm; uphold reusable chart and templating standards, namespaces, resource quotas, and Azure Policy for AKS.
• Strengthen AKS environments through policy enforcement, network policies, and image scanning.
• Oversee container and cluster vulnerability management, including scanning, triage, prioritization, and coordination of remediation efforts.
• Contribute to Terraform-based infrastructure as code for the provisioning and management of Azure resources.
• Support Azure DevOps or equivalent CI/CD pipelines, including GitOps workflows utilizing Flux/ArgoCD.
• Design, implement, and manage observability solutions across Grafana, Prometheus, Loki, Tempo, Azure Monitor, Log Analytics Workspace, and Application Insights.
• Define SLIs/SLOs and fine-tune alerting mechanisms to minimize noise.
• Identify, diagnose, and resolve performance bottlenecks across the AKS platform and its dependent services.
• Contribute to Disaster Recovery and Business Continuity Planning protocols, including failover drills and RTO/RPO validation.
• Provide escalation support for production Kubernetes and infrastructure incidents; engage in on-call rotation, lead root-cause analysis, and implement preventative measures.
• Document runbooks, post-incident reviews, and operational knowledge; maintain an evolving knowledge base.
• Over 5 years of practical Kubernetes experience.
• Strong understanding of Kubernetes internals, including scheduling, networking, storage, and RBAC.
• Expertise in Azure CNI networking and AKS private cluster configuration.
• Hands-on experience with Azure PaaS services such as ACR, AKV, Azure SQL, Kafka, and Managed Identities.
• Experience in troubleshooting network-layer dependencies, including DNS, firewall, and private endpoints.
• Proficient with Helm.
• Familiarity with Terraform for infrastructure as code.
• Working knowledge of Azure DevOps or similar CI/CD tools.
• Practical experience in container/cluster vulnerability management and remediation workflows.
• Scripting abilities in Bash and Python.
• Exceptional written and verbal communication skills; capable of clearly conveying technical issues to both technical and non-technical audiences.
• Willingness to participate in an on-call rotation may be required.
• Remote Work Environment.
• Flexible Time Away From Work Policy, including PTO, Personal, and Sick Days.
• Competitive Salary and Health/Medical Benefits.
• RRSP/TFSA/401K Employee Contribution.
• Life and Disability Insurance.
• Employee Assistance Program.
• FHIR Study Program and Skillsoft Learning.
• Super HAPI Fun Club.
• Workplace accommodations during interviews or while working at Smile.
Shift5
Exemplar Companies, PBC
Oddball
Experian
Get handpicked remote jobs straight to your inbox weekly.