
Site Reliability Engineer
Posted 1 day ago

Posted 1 day ago
This is a fully remote position, open to applicants in California.
• Enhance the availability, performance, scalability, and recoverability of AXON Networks' cloud solutions.
• Integrate software engineering with hands-on NOC operations to ensure the entire cloud-to-device service path is observable, supportable, and resilient at fleet scale.
• Develop practical SRE capabilities within the NOC while collaborating with Support, Operations, cloud, and DevOps Engineering.
• Take ownership of reliability outcomes for designated cloud services.
• Strengthen observability, capacity, resilience, and recovery measures.
• Define and implement service-level indicators, service-level objectives, and actionable alerting mechanisms.
• Automate repetitive NOC tasks and establish safe, testable processes for diagnosis, recovery, device operations, and routine production changes.
• Lead technical efforts during incidents, promote evidence-based learning, and ensure completion of corrective actions.
• Set reliability baselines, SLIs, SLOs, and error budgets for cloud services and device-management workflows.
• Trace failures across cloud APIs, microservices, Kubernetes, infrastructure, databases, messaging systems, networks, device-management protocols, and devices.
• Identify failure patterns that are fleet-wide and customer-specific.
• Provide operability requirements and production evidence during design and readiness reviews.
• Maintain NOC dashboards for tracking service health, device reachability, provisioning, command and telemetry performance, firmware adoption, and customer impact.
• Participate in the NOC production on-call rotation and act as the technical incident lead or senior troubleshooter.
• Diagnose complex failures in applications, infrastructure, Kubernetes, APIs, networking, databases, messaging systems, and CPE.
• Coordinate evidence gathering and technical escalation with customers, Engineering, firmware, DevOps, and vendors.
• Lead or assist in post-incident reviews and prioritize actionable corrective measures.
• Develop production-grade software, scripts, and workflows for diagnosis, remediation, deployment safety, fleet analysis, scaling, maintenance, and recovery.
• Enhance CI/CD and GitOps practices, including automated testing, release validation, progressive delivery, and rollback readiness.
• Manage or contribute to infrastructure as code, configuration as code, and reusable self-service patterns.
• Measure NOC toil and prioritize sustainable platform capabilities.
• Create capacity models for service-provider growth, managed devices, telemetry, messaging, API demand, and rollout events.
• Develop and maintain runbooks, troubleshooting decision trees, service maps, dependency records, and operational knowledge.
• Mentor NOC and Support personnel on diagnosis, mitigation, evidence capture, and escalation processes.
• Build self-service diagnostic tools and views to determine scope, affected customers, device cohorts, fault domains, and next actions.
• Share insights on reliability with Engineering and Product teams and contribute to reliability and operational-readiness reviews.
• Over 5 years of experience in site reliability engineering, production engineering, DevOps, cloud infrastructure, systems engineering, or a closely related field.
• Proficient in software development or automation using Python, Go, Java, Bash, or similar languages, with a track record of producing maintainable operational code.
• Direct experience managing distributed production systems in a public cloud environment and troubleshooting across application, infrastructure, network, and device-integration layers.
• Familiarity with Google Cloud Platform, Oracle Cloud Infrastructure, and production Kubernetes environments.
• Experience with infrastructure as code and delivery tools such as Terraform, Helm, Git-based CI/CD, and policy-as-code.
• Strong foundational knowledge of Linux, containers, and Kubernetes, including deployment behavior, resource management, networking, and failure diagnosis.
• Excellent troubleshooting and debugging capabilities within Kubernetes platforms.
• Experience with modern observability practices and tools, including metrics, logs, traces, alerting, dashboards, and synthetic monitoring.
• Knowledge of Prometheus, Grafana, OpenTelemetry, or equivalent observability ecosystems.
• Familiarity with Apache Pulsar or similar distributed messaging and streaming platforms that handle requests from millions of devices.
• Proven experience participating in an on-call rotation and responding effectively to high-severity, customer-impacting production incidents.
• Working knowledge of SLOs, error budgets, capacity planning, resilience engineering, change safety, and blameless incident analysis.
• Strong networking expertise, including TCP/IP, DNS, DHCP, TLS, routing, NAT, load balancing, and systematic troubleshooting at the packet or session level.
• Exceptional communication skills, disciplined documentation, and the ability to collaborate across NOC, cloud, DevOps, firmware, and service-provider teams.
• Bachelor’s degree in computer science, engineering, or equivalent practical experience.
• Preferred: Experience supporting cloud-managed CPEs, broadband gateways, routers, ONTs, Wi-Fi/mesh systems, or similar edge devices.
• Preferred: Familiarity with TR-069/CWMP, TR-369/USP, TR-181 data models, ACS or USP controller platforms, device telemetry, and remote lifecycle management.
• Preferred: Experience supporting Apache Pulsar or Kafka, APIs, and highly available databases utilized in device-management control planes.
• Preferred: Understanding of GPON/XGS-PON, DOCSIS, Ethernet, or fixed wireless access technologies.
• Preferred: Experience with firmware rollout automation, canary or cohort deployments, fleet health analysis, and safe rollback practices.
• Preferred: Experience building auto-remediation, safe self-service operations, or internal reliability platforms.
• Preferred: Experience supporting multiple service-provider customers in a 24/7 telecommunications, broadband, or managed-network setting.
• Must be eligible to work without visa sponsorship.
• Competitive salary and performance-based bonuses.
• Comprehensive health, dental, and vision insurance.
• Opportunities for professional development and continuous learning.
• Flexible working hours and remote work options.
• Generous paid time off and holiday leave.
• Collaborative and inclusive company culture.
Mirantis
HumanIT Digital Consulting
Gormat
Wizeline
Get handpicked remote jobs straight to your inbox weekly.