
Site Reliability Engineer
Posted 1 day ago

Posted 1 day ago
This is a fully remote position, open to applicants in Canada.
• Enhance the availability, performance, scalability, and recoverability of AXON Networks' cloud solutions.
• Integrate software engineering with hands-on NOC operations to ensure the cloud-to-device service path is observable, supportable, and resilient at fleet scale.
• Establish effective SRE capabilities within the NOC while collaborating with Support, Operations, cloud, and DevOps Engineering teams.
• Engage in a sustainable on-call rotation and enhance the diagnosis of customer-impacting issues.
• Take ownership of reliability outcomes for designated cloud services.
• Improve observability, capacity, resilience, and recovery processes.
• Define and operationalize SLIs, SLOs, and actionable alerting mechanisms.
• Automate repetitive NOC tasks and develop safe, testable procedures for diagnosis, recovery, device operations, and routine production changes.
• Lead technical efforts during incidents, promote evidence-based learning, and implement corrective actions.
• Establish reliability baselines, SLIs, SLOs, and error budgets for cloud services and device management workflows.
• Trace failures across cloud APIs, microservices, Kubernetes, infrastructure, databases, messaging systems, networks, device-management protocols, and devices.
• Identify failure patterns that affect both fleet-wide and specific customer operations.
• Contribute operability requirements and production evidence during design and readiness reviews.
• Maintain NOC dashboards that monitor service health, device reachability, provisioning, commands, telemetry, firmware adoption, and customer impact.
• Participate in the NOC production on-call rotation as a technical incident lead or senior troubleshooter.
• Diagnose complex failures across applications, cloud infrastructure, Kubernetes, APIs, networking, DNS/TLS, databases, messaging, device-management sessions, and CPE behavior.
• Coordinate evidence gathering and technical escalation with customers, Engineering, firmware, DevOps, and vendors.
• Lead or assist in post-incident reviews and transform recurring failures into measurable corrective actions.
• Develop production-grade software, scripts, and workflows for diagnosis, remediation, deployment safety, fleet analysis, scaling, maintenance, and recovery.
• Enhance CI/CD and GitOps practices, including automated testing, release validation, progressive delivery, and rollback readiness.
• Manage or contribute to infrastructure as code, configuration as code, and reusable self-service patterns.
• Measure NOC toil and prioritize sustainable platform capabilities with Automation & Tools Engineers.
• Create capacity models for service-provider growth, managed-device populations, telemetry, messaging, APIs, and rollout events.
• Develop and maintain runbooks, troubleshooting decision trees, service maps, dependency records, known-error guidance, and operational knowledge.
• Mentor NOC and Support personnel on diagnosis, mitigation, evidence capture, and escalation processes.
• Build self-service diagnostic tools and views to ascertain scope, affected customers, device cohorts, fault domains, and next actions.
• Share reliability insights with Engineering and Product teams, contributing to reliability and operational-readiness reviews.
• More than 5 years of experience in site reliability engineering, production engineering, DevOps, cloud infrastructure, systems engineering, or a closely related field.
• Strong software or automation skills in Python, Go, Java, Bash, or a similar language, with a proven track record of producing maintainable operational code.
• Hands-on experience managing distributed production systems in a public cloud environment and troubleshooting across application, infrastructure, network, and device-integration layers.
• Familiarity with Google Cloud Platform, Oracle Cloud Infrastructure, and production Kubernetes environments.
• Experience with Terraform, Helm, Git-based CI/CD, and policy-as-code implementations.
• Strong foundational knowledge in Linux, containers, and Kubernetes, including deployment behavior, resource management, networking, and failure diagnosis.
• Proven troubleshooting and debugging skills within Kubernetes platforms.
• Experience with observability practices and tools covering metrics, logs, traces, alerting, dashboards, and synthetic monitoring.
• Familiarity with Prometheus, Grafana, OpenTelemetry, or similar observability ecosystems.
• Knowledge of Apache Pulsar or comparable distributed messaging and streaming platforms managing requests from millions of devices.
• Experience participating in an on-call rotation and responding to high-severity, customer-impacting production incidents.
• Working knowledge of SLOs, error budgets, capacity planning, resilience engineering, change safety, and blameless incident learning methodologies.
• Strong networking knowledge, including TCP/IP, DNS, DHCP, TLS, routing, NAT, load balancing, and troubleshooting at the packet or session level.
• Excellent communication skills, meticulous documentation practices, and the ability to collaborate across NOC, cloud, DevOps, firmware, and service-provider teams.
• Bachelor’s degree in computer science, engineering, or equivalent practical experience.
• Preferred: Experience supporting cloud-managed CPEs in a service-provider setting.
• Preferred: Familiarity with TR-069/CWMP, TR-369/USP, TR-181 data models, ACS or USP controller platforms, device telemetry, and remote lifecycle management.
• Preferred: Experience with Apache Pulsar or Kafka, APIs, and highly available databases used in device management control planes.
• Preferred: Understanding of GPON/XGS-PON, DOCSIS, Ethernet, or fixed wireless access technologies.
• Preferred: Experience with firmware rollout automation, canary or cohort deployments, fleet health analysis, and safe rollback practices.
• Preferred: Experience building auto-remediation, safe self-service operations, or internal reliability platforms.
• Preferred: Experience supporting multiple service-provider customers in a 24×7 telecommunications, broadband, or managed-network environment.
• Equal opportunity recruitment process.
• Inclusive and diverse working environment.
Mirantis
HumanIT Digital Consulting
Gormat
Wizeline
Get handpicked remote jobs straight to your inbox weekly.