Remotery

Customer Reliability Engineer

Posted Aug 4

This is a fully remote position, open to applicants in California, +4 more states.

📋 Description

• Take ownership of Cisco Hypershield cases escalated from Cisco TAC until resolution, engaging with customers directly as necessary.

• Diagnose intricate production failures on the N9300 Smart Switch fabric and within on-premises Kubernetes controllers.

• Identify faults across switching, forwarding, security services and enforcement, as well as control-plane layers.

• Comprehend customer architectures and configurations to diagnose issues in unfamiliar production settings.

• Replicate customer failures, collaborate with engineering to implement fixes, and ensure the resolution is communicated back to the customer.

• Transform individual cases into systemic enhancements, such as runbooks, diagnostics, knowledge-base content, and product feedback.

• Develop the team’s proactive perspective on customer health through monitoring, tools, and reliability practices.


⛳️ Requirements

• Bachelor’s degree plus 8 years of experience, Master’s degree plus 6 years, or equivalent industry experience.

• Proven experience supporting enterprise customers in an escalation role.

• Demonstrated ability to diagnose and resolve complex production incidents under SLA pressures in unfamiliar environments.

• Experience operating and troubleshooting Cisco Nexus / NX-OS in production; equivalent expertise on another major vendor is acceptable.

• Skill in localizing failures across layered data-center architectures that include switching/forwarding, services/enforcement, and control-plane domains.

• Proficient in Linux operations at the command line, including troubleshooting in production environments.

• Working knowledge of containers or Kubernetes.

• Familiarity with packet capture and flow-telemetry analysis, including NetFlow/IPFIX.

• Understanding of enterprise virtualization and troubleshooting VM-based appliance deployments; vSphere is the current deployment target.

• Proficiency in operational Kubernetes and Helm, including TLS certificates, service-account authentication, API-server connectivity, service exposure, persistent storage, custom resources, and operators.

• Knowledge of VXLAN EVPN fabrics.

• Understanding of network segmentation and firewall policy design, including zone-based or microsegmentation methods.

• Awareness of the NetOps/NetSecOps operational split in data-center security.

• Experience guiding diagnosis and remediation through a customer’s own team in environments without direct access.

• Familiarity with NX-OS automation and APIs such as NX-API, NETCONF/RESTCONF, gNMI, or Ansible.

• Ability to effectively communicate incident status, root causes, and remediation steps to both technical and executive audiences, both verbally and in writing.

• CCNP Data Center, CCNP Enterprise, DevNet Professional, CCIE Data Center, CCIE Enterprise, or DevNet Expert certifications are advantageous, but not mandatory.


🏝️ Benefits

• Medical, dental, and vision insurance.

• 401(k) plan with Cisco matching contributions.

• Paid parental leave.

• Short- and long-term disability coverage.

• Basic life insurance.

• Restricted stock unit grants may be available, contingent on continued employment and vesting periods.

• 10 paid holidays per full calendar year.

• 1 floating holiday for non-exempt employees.

• 1 paid day off on the employee’s birthday.

• Paid year-end holiday shutdown.

• 4 paid days off for personal wellness.

• 16 days of paid vacation time per full calendar year for non-exempt employees.

• Flexible vacation time off program with no defined limit for eligible exempt employees.

• 80 hours of sick time off provided upon hire and each January 1st thereafter.

• Up to 80 hours of unused sick time may be carried forward.

• Additional paid time off for critical or emergency family issues.

• Optional 10 paid volunteer days per full calendar year.

• Annual bonuses may be available for non-sales roles.

• Opportunities for growth and development on a global scale.

People also viewed

DATAGROUP1 day ago

DevOps Engineer

DE flagGermany OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Ambush1 day ago

DevOps Engineer

BR flagBrazil OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
DuoKey1 day ago

DevOps Engineer

MU flagMauritius OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
TEKsystems2 days ago

SRE – CloudOps, Practice Architect II

US flagIllinois OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$148.2k – $222.4k/year
ApplyView job
TEKsystems2 days ago

SRE CloudOps Practice Architect II

US flagTexas OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$148.2k – $222.4k/year
ApplyView job
Level Data2 days ago

Senior DevOps Engineer

US flagMassachusetts OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$120k – $135k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers