
Senior Network Reliability Engineer
Posted 6 days ago

Posted 6 days ago
This is a fully remote position, open to applicants in United States.
• Establish network-platform SLOs and error budgets to regulate changes.
• Facilitate postmortems aimed at long-term remediation.
• Transition network state into declarative, version-controlled code utilizing Terraform or Pulumi, Ansible, Python, and Infra CI/CD.
• Develop network policy as intent and enforce Policy as Code controls.
• Design, implement, and oversee network infrastructure through Infrastructure as Code.
• Operate and enhance the multi-account AWS Landing Zone, encompassing Cloud WAN, Transit Gateway, IPAM, private DNS, and telemetry pipelines.
• Create platform abstractions for the proper onboarding of accounts and services.
• Engineer Kubernetes networking, service mesh, eBPF observability, and cloud-tier identity integrations.
• Establish structured alerts, runbooks, and monitoring dashboards for both operators and end-users.
• Offer technical guidance to junior engineers and mentor the team.
• Maintain and configure routers, switches, and firewalls at data centers and offices as needed.
• Act as an escalation point for network incidents and conduct troubleshooting through root-cause analysis.
• Write runbooks and SOPs, packaging routine work for L1/L2 handoff.
• Coordinate reliability practices across Data Platforms, NOC/SOC, and Cyber Security.
• In-depth knowledge of TCP/IP, BGP, OSPF, VPNs, and SD-WAN architecture.
• Solid production experience with Terraform state management and modules, Ansible playbooks and roles, or equivalent.
• Proficient in Python for automation and API interaction, or similar languages.
• Practical experience with Cloudflare, Zscaler, and/or enterprise firewalls.
• Experience in configuring monitoring tools such as Datadog, Prometheus, or Grafana for effective alerts and dashboards.
• Familiarity with service mesh technologies like Istio, Linkerd, Consul Connect, or Cilium (a plus).
• Experience with eBPF-based observability tools such as Hubble or Pixie (a plus).
• Knowledge of AWS multi-account landing-zone tooling, including AFT or Control Tower (a plus).
• Experience with Policy as Code using OPA/Rego, Sentinel, or Cilium NetworkPolicy (a plus).
• Strong commitment to documentation-first practices.
• A mindset geared towards automating repetitive tasks and minimizing toil.
• Openness to handling physical hardware tasks when necessary while maintaining a software-centric engineering approach.
• Comprehensive health, dental, and vision insurance plan options.
• Basic and Supplemental Life Insurance.
• Short- and Long-Term Disability coverage.
• Employee Assistance Program.
• Wellness programs.
• 401K plan with matching contributions from the Company.
• Supportive work environment that values employee diversity.
DATAGROUP
Ambush
DuoKey
TEKsystems
Get handpicked remote jobs straight to your inbox weekly.