
Network Automation, Reliability Engineer
Posted Aug 7

Posted Aug 7
This is a fully remote position, open to applicants in California, +2 more states.
• Design, develop, and maintain Python tools that provision, validate, and audit network devices across the global infrastructure.
• Create model-driven provisioning libraries utilizing Jinja2 templating with NETCONF/YANG or REST APIs to codify optimal configurations and prevent configuration drift.
• Develop automated solutions for recurring failure categories triggered by syslog and telemetry.
• Identify repetitive manual runbook processes and replace them with thoroughly tested and reviewed code.
• Manage and scale data center fabrics, backbone links, and out-of-band management networks across data centers and POP locations.
• Design and optimize BGP policies to control path selection and eliminate single-carrier points of failure.
• Assist in capacity expansion and site turn-ups, encompassing high-level design, port channelization, optics and cabling standards, and hardware validation.
• Implement zero-downtime production change management, including hitless migrations and staged rollbacks.
• Operationalize multi-vendor streaming telemetry and fine-tune subscription jobs.
• Build and sustain observability for hop-by-hop path tracing, multi-layer fault isolation, and root cause analysis.
• Contribute to a real-time network source of truth that aggregates BGP, link-state, and drain-state data.
• Participate in a 24x7 on-call rotation, lead incident response, document RCAs, and drive follow-up automation.
• Maintain and enhance firewall and ACL policies across multi-vendor platforms.
• Support SIRT/PSIRT CVE remediation through automated regression testing and configuration-as-code pipelines.
• Create and maintain technical documentation, including designs, runbooks, and API agreements.
• Proven, sustained Python development experience in a production network or infrastructure setting.
• Experience with production tooling for configuration generation and validation, API integrations, telemetry collectors, automated remediation, or test harnesses.
• Proficient with modules, packaging, testing, code reviews, and version control.
• Hands-on production BGP experience, including policy, path selection, and multihoming.
• Familiarity with IS-IS or OSPF, ECMP, and VXLAN/EVPN or MPLS overlays.
• Operational experience with at least two of Arista EOS, Juniper Junos (QFX/SRX/PTX/MX), or Cisco IOS-XR/NX-OS.
• Experience with Ansible and Jinja2.
• Knowledge of NETCONF/YANG, RESTCONF, or vendor REST APIs.
• Familiarity with gNMI/gRPC streaming telemetry, OpenConfig models, SNMP, flow telemetry, and dashboarding/alerting.
• Linux experience, including networking stack, packet capture, systemd services, and shell scripting.
• Git-based workflows with peer review and CI pipelines such as Jenkins, GitLab CI, or GitHub Actions.
• Experience participating in a 24x7 production on-call rotation, including incident command and root cause analysis.
• Approximately 2–5 years of experience in network production, network reliability, or network automation.
• Must reside in the United States and be authorized to work in the US.
• Preferred: out-of-band network experience, data center or POP build-out, MACsec/IPsec, 802.1X/NAC, optical or transport technologies, NetBox or in-house source of truth, LLM or agentic on-call tooling, and a master's degree.
• Competitive salary and performance-based bonuses.
• Comprehensive health, dental, and vision insurance.
• Generous paid time off and holiday schedule.
• Opportunities for professional development and continuing education.
• Flexible work arrangements and remote work options.
CVS Health
Devoteam
Aspirion
Goodgame Studios
Get handpicked remote jobs straight to your inbox weekly.