
Network Automation Lead
Posted Aug 7

Posted Aug 7
This is a fully remote position, open to applicants in Nevada.
• Take ownership of the complete zero-touch provisioning pipeline, guiding it from bare-metal switch boot to production readiness without any human intervention.
• Develop intent-based configuration generation utilizing network source-of-truth/IPAM data, alongside GitOps deployment, validation, rollout, and rollback processes.
• Implement network validation and pre-deployment snapshot/digital-twin testing protocols.
• Create gNMI/OpenConfig streaming telemetry, metrics, and logging pipelines to assess fabric health.
• Monitor RoCE health, PFC/ECN counters, optics and link errors, BGP/EVPN state, capacity, and utilization.
• Provide dashboards and alerting solutions for the network team.
• Collect requirements to develop self-service APIs and user interfaces.
• Collaborate with Network Engineering to ensure that tools align with network operations and turn-up processes.
• Recruit, mentor, and develop a small team of software engineers and SREs; manage roadmap, prioritization, and delivery while also contributing code.
• Establish technical direction and standards, integrating with infrastructure provisioning, CI/CD, secrets, identity, and observability tools.
• Employ software engineering best practices through code reviews, testing, release management, and on-call responsibilities.
• Over 8 years of pertinent experience.
• Demonstrated capability in building network automation at scale, ideally within a hyperscaler, large cloud, or extensive datacenter/AI-infrastructure operator.
• Experience as a core contributor to an end-to-end ZTP/device-provisioning system.
• Strong software engineering fundamentals in Python and/or Go, including version control, testing, CI/CD, and code review.
• Hands-on expertise with datacenter Clos fabrics and BGP, EVPN/VXLAN, and preferably RoCEv2/RDMA.
• Proficiency in gNMI/gNOI, OpenConfig/YANG, NETCONF, and source-of-truth systems like NetBox/Nautobot, as well as NOS platforms such as SONiC/FRR or their equivalents, and tools like Nornir/NAPALM/Ansible.
• Experience with observability/telemetry pipelines such as Prometheus, Grafana, Kafka, or OpenTelemetry.
• Comfortable operating services on Kubernetes/containers.
• Leadership experience: either led a team or acted as the clear technical owner of a platform.
• Preferred: Experience with GPU/AI training or inference clusters.
• Preferred: Familiarity with the AMD networking ecosystem, Pensando DPUs, Ultra Ethernet, or Ethernet-based RDMA fabrics.
• Preferred: Knowledge of whitebox/disaggregated networking and SONiC at scale.
• Preferred: Experience with network validation/digital-twin tools such as Batfish or containerlab.
• Preferred: Involvement in multi-site/multi-region datacenter buildouts.
• Authorization to work in the United States is required.
• Stock Options
• 100% paid Medical, Dental, and Vision insurance for Employees
• Company Health Savings Account Contributions
• 100% paid Short Term and Long Term Disability Insurance for Employees
• Life and Voluntary Supplemental Insurance Options
• Other Insurance Options, such as Pet & Legal Insurance
• Various Supplementary Health Benefits, including discounted Virtual Healthcare Appointments and Serious Illness Support
• Flexible Spending Account
• 401(k)
• Employee Assistance Program
• Flexible PTO
• Paid Holidays
• Parental Leave
• Other In-Office Perks
Cloudera
Stellar Cyber
Pragmatike
Pragmatike
Get handpicked remote jobs straight to your inbox weekly.