Site Reliability Engineering Technical Leader

Posted 4 days ago

This is a fully remote position, open to applicants in California, +4 more states.

📋 Description

• Oversee the feasibility evaluation, technical planning, and phased transition of suitable workloads from AWS-hosted Kubernetes clusters to the internal Kubernetes platform.

• Define technical direction and develop staged delivery strategies.

• Aid in the operation and reliability of specialized Kubernetes workloads that have unique load-balancing, networking, and traffic requirements.

• Provide support for applications post-migration, including ensuring production readiness, incident response, performance enhancement, and operational improvements.

• Tackle complex challenges across applications, Kubernetes, Linux, networking, containers, and infrastructure.

• Enhance the scalability, reliability, security, performance, and operability of services hosted on Kubernetes.

• Collaborate with Node Connectivity, firmware, cloud infrastructure, security, Site Reliability Engineering (SRE), and product teams.

• Prioritize support and manage cross-system modifications.

• Contribute to wider Kubernetes SRE and platform reliability projects.

• Design, implement, and maintain production-quality Go software for distributed, concurrent, and networked systems.

• Promote operational readiness through Service Level Indicators (SLIs)/Service Level Objectives (SLOs), thorough testing, on-call support, root-cause analysis, and long-term corrective measures.


⛳️ Requirements

• A minimum of 10 years of professional experience in software, site reliability, or infrastructure engineering, including technical leadership of significant production systems.

• Proven experience in designing, deploying, and managing large distributed services on Kubernetes.

• At least 5 years of programming experience in Go or a comparable systems programming language.

• Experience in supporting production services through incident response, performance analysis, Kubernetes reliability practices, observability, automation, and operational readiness.

• Strong knowledge of Linux and networking concepts, encompassing IPv4/IPv6, TCP, routing, DNS, and TLS.

• Excellent judgment in architecture, incident response, prioritization, and technical trade-offs.

• Proven ability to align stakeholders across teams without relying on formal authority.

• Preferred: Experience in operating Kubernetes in AWS, on-premises, or hybrid environments utilizing AWS/EKS, Docker, Kustomize, GitLab CI/CD, or similar deployment systems.

• Preferred: Familiarity with Kubernetes networking, ingress, load balancing, service discovery, and traffic management for high-throughput or geographically distributed services.

• Preferred: Experience with gRPC, Protocol Buffers, mutual TLS, PKI, VPNs, tunneling, or network security.

• Preferred: Experience with OpenTelemetry, Prometheus, Datadog, or similar observability tools, as well as distributed routing, packet processing, performance optimization, or failure testing.


🏝️ Benefits

• Medical, dental, and vision insurance.

• 401(k) plan with a Cisco matching contribution.

• Paid parental leave.

• Short- and long-term disability coverage.

• Basic life insurance.

• Cisco restricted stock unit grants may be available, subject to eligibility and continued employment.

• 10 paid holidays per full calendar year.

• 1 floating holiday for non-exempt employees.

• Paid birthday day off.

• Paid year-end holiday shutdown.

• 4 paid personal wellness days.

• 16 days of paid vacation per full calendar year for non-exempt employees.

• Flexible vacation time off program with no defined limit for eligible exempt employees.

• 80 hours of sick time granted on hire date and each January 1st thereafter.

• Up to 80 hours of unused sick time may be carried forward.

• Additional paid time off for critical or emergency family situations.

• Optional 10 paid volunteer days per full calendar year.

• Annual bonuses for non-sales roles, subject to Cisco policies.

• Incentive compensation for sales-plan employees, subject to applicable Cisco plan.

People also viewed

Koniag Government Services2 days ago

Architect/DevOps Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
FP Markets (First Prudential Markets)2 days ago

Senior DevOps Engineer

AM flagArmenia, +4 more countriesFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Modern Campus2 days ago

Senior DevOps Engineer

US flagUnited States, +1 more countryFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
InRule2 days ago

Site Reliability Engineer

US flagUnited States, +1 more countryFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Thumbtack2 days ago

Senior Software Engineer, Site Reliability Engineering

US flagUnited States, +38 more locationsFull-timeDevOps & Site Reliability Engineer (SRE)$179.4k – $272.8k/year
ApplyView job
Thumbtack2 days ago

Senior Software Engineer, Site Reliability Engineering

CA flagCanada, +1 more countryFull-timeDevOps & Site Reliability Engineer (SRE)C$180.2k – C$233.2k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers