Site Reliability Engineer

atAXON NetworksRemoteCA flagCanadaFreelanceDevOps & Site Reliability Engineer (SRE)Mid-levelSeniorC$80 – C$110/hour

Posted 1 day ago

This is a fully remote position, open to applicants in Canada.

📋 Description

• Enhance the availability, performance, scalability, and recoverability of AXON Networks' cloud solutions.

• Integrate software engineering with hands-on NOC operations to ensure the cloud-to-device service path is observable, supportable, and resilient at fleet scale.

• Establish effective SRE capabilities within the NOC while collaborating with Support, Operations, cloud, and DevOps Engineering teams.

• Engage in a sustainable on-call rotation and enhance the diagnosis of customer-impacting issues.

• Take ownership of reliability outcomes for designated cloud services.

• Improve observability, capacity, resilience, and recovery processes.

• Define and operationalize SLIs, SLOs, and actionable alerting mechanisms.

• Automate repetitive NOC tasks and develop safe, testable procedures for diagnosis, recovery, device operations, and routine production changes.

• Lead technical efforts during incidents, promote evidence-based learning, and implement corrective actions.

• Establish reliability baselines, SLIs, SLOs, and error budgets for cloud services and device management workflows.

• Trace failures across cloud APIs, microservices, Kubernetes, infrastructure, databases, messaging systems, networks, device-management protocols, and devices.

• Identify failure patterns that affect both fleet-wide and specific customer operations.

• Contribute operability requirements and production evidence during design and readiness reviews.

• Maintain NOC dashboards that monitor service health, device reachability, provisioning, commands, telemetry, firmware adoption, and customer impact.

• Participate in the NOC production on-call rotation as a technical incident lead or senior troubleshooter.

• Diagnose complex failures across applications, cloud infrastructure, Kubernetes, APIs, networking, DNS/TLS, databases, messaging, device-management sessions, and CPE behavior.

• Coordinate evidence gathering and technical escalation with customers, Engineering, firmware, DevOps, and vendors.

• Lead or assist in post-incident reviews and transform recurring failures into measurable corrective actions.

• Develop production-grade software, scripts, and workflows for diagnosis, remediation, deployment safety, fleet analysis, scaling, maintenance, and recovery.

• Enhance CI/CD and GitOps practices, including automated testing, release validation, progressive delivery, and rollback readiness.

• Manage or contribute to infrastructure as code, configuration as code, and reusable self-service patterns.

• Measure NOC toil and prioritize sustainable platform capabilities with Automation & Tools Engineers.

• Create capacity models for service-provider growth, managed-device populations, telemetry, messaging, APIs, and rollout events.

• Develop and maintain runbooks, troubleshooting decision trees, service maps, dependency records, known-error guidance, and operational knowledge.

• Mentor NOC and Support personnel on diagnosis, mitigation, evidence capture, and escalation processes.

• Build self-service diagnostic tools and views to ascertain scope, affected customers, device cohorts, fault domains, and next actions.

• Share reliability insights with Engineering and Product teams, contributing to reliability and operational-readiness reviews.


⛳️ Requirements

• More than 5 years of experience in site reliability engineering, production engineering, DevOps, cloud infrastructure, systems engineering, or a closely related field.

• Strong software or automation skills in Python, Go, Java, Bash, or a similar language, with a proven track record of producing maintainable operational code.

• Hands-on experience managing distributed production systems in a public cloud environment and troubleshooting across application, infrastructure, network, and device-integration layers.

• Familiarity with Google Cloud Platform, Oracle Cloud Infrastructure, and production Kubernetes environments.

• Experience with Terraform, Helm, Git-based CI/CD, and policy-as-code implementations.

• Strong foundational knowledge in Linux, containers, and Kubernetes, including deployment behavior, resource management, networking, and failure diagnosis.

• Proven troubleshooting and debugging skills within Kubernetes platforms.

• Experience with observability practices and tools covering metrics, logs, traces, alerting, dashboards, and synthetic monitoring.

• Familiarity with Prometheus, Grafana, OpenTelemetry, or similar observability ecosystems.

• Knowledge of Apache Pulsar or comparable distributed messaging and streaming platforms managing requests from millions of devices.

• Experience participating in an on-call rotation and responding to high-severity, customer-impacting production incidents.

• Working knowledge of SLOs, error budgets, capacity planning, resilience engineering, change safety, and blameless incident learning methodologies.

• Strong networking knowledge, including TCP/IP, DNS, DHCP, TLS, routing, NAT, load balancing, and troubleshooting at the packet or session level.

• Excellent communication skills, meticulous documentation practices, and the ability to collaborate across NOC, cloud, DevOps, firmware, and service-provider teams.

• Bachelor’s degree in computer science, engineering, or equivalent practical experience.

• Preferred: Experience supporting cloud-managed CPEs in a service-provider setting.

• Preferred: Familiarity with TR-069/CWMP, TR-369/USP, TR-181 data models, ACS or USP controller platforms, device telemetry, and remote lifecycle management.

• Preferred: Experience with Apache Pulsar or Kafka, APIs, and highly available databases used in device management control planes.

• Preferred: Understanding of GPON/XGS-PON, DOCSIS, Ethernet, or fixed wireless access technologies.

• Preferred: Experience with firmware rollout automation, canary or cohort deployments, fleet health analysis, and safe rollback practices.

• Preferred: Experience building auto-remediation, safe self-service operations, or internal reliability platforms.

• Preferred: Experience supporting multiple service-provider customers in a 24×7 telecommunications, broadband, or managed-network environment.


🏝️ Benefits

• Equal opportunity recruitment process.

• Inclusive and diverse working environment.

People also viewed

Mirantis23 hours ago

Senior Site Reliability Engineer, Golang, Kubernetes

KZ flagKazakhstan, +1 more countryFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
HumanIT Digital Consulting1 day ago

DevOps Engineer – AWS, Kubernetes, Terraform

PT flagPortugal OnlyPart-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Gormat1 day ago

Cloud DevOps Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$191k – $206k/year
ApplyView job
Wizeline1 day ago

Site Reliability Engineer – Cloud Security, Posture Management

AR flagArgentina OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Stefanini Brasil1 day ago

DevOps Specialist

BR flagBrazil OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Zigabyte1 day ago

DevSecOps Engineer

US flagVirginia OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers