Site Reliability Engineer

atAXON NetworksRemoteUS flagCaliforniaFull-timeDevOps & Site Reliability Engineer (SRE)Mid-levelSenior$160k – $200k/year

Posted 1 day ago

This is a fully remote position, open to applicants in California.

📋 Description

• Enhance the availability, performance, scalability, and recoverability of AXON Networks' cloud solutions.

• Integrate software engineering with hands-on NOC operations to ensure the entire cloud-to-device service path is observable, supportable, and resilient at fleet scale.

• Develop practical SRE capabilities within the NOC while collaborating with Support, Operations, cloud, and DevOps Engineering.

• Take ownership of reliability outcomes for designated cloud services.

• Strengthen observability, capacity, resilience, and recovery measures.

• Define and implement service-level indicators, service-level objectives, and actionable alerting mechanisms.

• Automate repetitive NOC tasks and establish safe, testable processes for diagnosis, recovery, device operations, and routine production changes.

• Lead technical efforts during incidents, promote evidence-based learning, and ensure completion of corrective actions.

• Set reliability baselines, SLIs, SLOs, and error budgets for cloud services and device-management workflows.

• Trace failures across cloud APIs, microservices, Kubernetes, infrastructure, databases, messaging systems, networks, device-management protocols, and devices.

• Identify failure patterns that are fleet-wide and customer-specific.

• Provide operability requirements and production evidence during design and readiness reviews.

• Maintain NOC dashboards for tracking service health, device reachability, provisioning, command and telemetry performance, firmware adoption, and customer impact.

• Participate in the NOC production on-call rotation and act as the technical incident lead or senior troubleshooter.

• Diagnose complex failures in applications, infrastructure, Kubernetes, APIs, networking, databases, messaging systems, and CPE.

• Coordinate evidence gathering and technical escalation with customers, Engineering, firmware, DevOps, and vendors.

• Lead or assist in post-incident reviews and prioritize actionable corrective measures.

• Develop production-grade software, scripts, and workflows for diagnosis, remediation, deployment safety, fleet analysis, scaling, maintenance, and recovery.

• Enhance CI/CD and GitOps practices, including automated testing, release validation, progressive delivery, and rollback readiness.

• Manage or contribute to infrastructure as code, configuration as code, and reusable self-service patterns.

• Measure NOC toil and prioritize sustainable platform capabilities.

• Create capacity models for service-provider growth, managed devices, telemetry, messaging, API demand, and rollout events.

• Develop and maintain runbooks, troubleshooting decision trees, service maps, dependency records, and operational knowledge.

• Mentor NOC and Support personnel on diagnosis, mitigation, evidence capture, and escalation processes.

• Build self-service diagnostic tools and views to determine scope, affected customers, device cohorts, fault domains, and next actions.

• Share insights on reliability with Engineering and Product teams and contribute to reliability and operational-readiness reviews.


⛳️ Requirements

• Over 5 years of experience in site reliability engineering, production engineering, DevOps, cloud infrastructure, systems engineering, or a closely related field.

• Proficient in software development or automation using Python, Go, Java, Bash, or similar languages, with a track record of producing maintainable operational code.

• Direct experience managing distributed production systems in a public cloud environment and troubleshooting across application, infrastructure, network, and device-integration layers.

• Familiarity with Google Cloud Platform, Oracle Cloud Infrastructure, and production Kubernetes environments.

• Experience with infrastructure as code and delivery tools such as Terraform, Helm, Git-based CI/CD, and policy-as-code.

• Strong foundational knowledge of Linux, containers, and Kubernetes, including deployment behavior, resource management, networking, and failure diagnosis.

• Excellent troubleshooting and debugging capabilities within Kubernetes platforms.

• Experience with modern observability practices and tools, including metrics, logs, traces, alerting, dashboards, and synthetic monitoring.

• Knowledge of Prometheus, Grafana, OpenTelemetry, or equivalent observability ecosystems.

• Familiarity with Apache Pulsar or similar distributed messaging and streaming platforms that handle requests from millions of devices.

• Proven experience participating in an on-call rotation and responding effectively to high-severity, customer-impacting production incidents.

• Working knowledge of SLOs, error budgets, capacity planning, resilience engineering, change safety, and blameless incident analysis.

• Strong networking expertise, including TCP/IP, DNS, DHCP, TLS, routing, NAT, load balancing, and systematic troubleshooting at the packet or session level.

• Exceptional communication skills, disciplined documentation, and the ability to collaborate across NOC, cloud, DevOps, firmware, and service-provider teams.

• Bachelor’s degree in computer science, engineering, or equivalent practical experience.

• Preferred: Experience supporting cloud-managed CPEs, broadband gateways, routers, ONTs, Wi-Fi/mesh systems, or similar edge devices.

• Preferred: Familiarity with TR-069/CWMP, TR-369/USP, TR-181 data models, ACS or USP controller platforms, device telemetry, and remote lifecycle management.

• Preferred: Experience supporting Apache Pulsar or Kafka, APIs, and highly available databases utilized in device-management control planes.

• Preferred: Understanding of GPON/XGS-PON, DOCSIS, Ethernet, or fixed wireless access technologies.

• Preferred: Experience with firmware rollout automation, canary or cohort deployments, fleet health analysis, and safe rollback practices.

• Preferred: Experience building auto-remediation, safe self-service operations, or internal reliability platforms.

• Preferred: Experience supporting multiple service-provider customers in a 24/7 telecommunications, broadband, or managed-network setting.

• Must be eligible to work without visa sponsorship.


🏝️ Benefits

• Competitive salary and performance-based bonuses.

• Comprehensive health, dental, and vision insurance.

• Opportunities for professional development and continuous learning.

• Flexible working hours and remote work options.

• Generous paid time off and holiday leave.

• Collaborative and inclusive company culture.

People also viewed

Mirantis23 hours ago

Senior Site Reliability Engineer, Golang, Kubernetes

KZ flagKazakhstan, +1 more countryFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
HumanIT Digital Consulting1 day ago

DevOps Engineer – AWS, Kubernetes, Terraform

PT flagPortugal OnlyPart-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Gormat1 day ago

Cloud DevOps Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$191k – $206k/year
ApplyView job
Wizeline1 day ago

Site Reliability Engineer – Cloud Security, Posture Management

AR flagArgentina OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Stefanini Brasil1 day ago

DevOps Specialist

BR flagBrazil OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Zigabyte1 day ago

DevSecOps Engineer

US flagVirginia OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers