
Site Reliability Engineering Lead
Posted 3 days ago

Posted 3 days ago
This is a fully remote position, open to applicants in Connecticut, +5 more states.
• Oversee, mentor, and develop a team of Site Reliability Engineers (SREs); conduct one-on-one meetings, performance evaluations, and career development strategies.
• Take charge of recruitment, onboarding, and decisions regarding team capacity and resources.
• Establish team objectives, prioritize tasks in the backlog, and facilitate planning sessions.
• Promote a culture of blameless post-incident analysis and encourage collaboration across Development, Security, and Product teams.
• Lead initiatives aimed at enhancing reliability throughout infrastructure and services.
• Manage incident response activities and focus on continual service improvement.
• Advocate for automation and operational excellence throughout the platform.
• Assist in developing scalable, secure, and resilient cloud-native environments.
• Provide line management for a small to medium-sized team, overseeing performance management, compensation, and recruitment authority.
• Facilitate post-mortem reviews and ensure timely creation of Root Cause Analyses (RCAs).
• Extensive knowledge of Kubernetes, including its architecture, upgrades, autoscaling, security hardening, and troubleshooting at scale.
• Proficient experience with Terraform, covering modular Infrastructure as Code (IaC) design, state management, multi-environment provisioning, and policy-as-code.
• In-depth understanding of Azure Cloud, encompassing compute, networking, identity (Azure Active Directory), storage, and cost optimization.
• Experience in designing and scaling CI/CD pipelines utilizing GitHub Actions, including release strategies and rollback automation.
• Familiarity with observability platforms such as Prometheus, Grafana, OpenTelemetry, and management of SLOs, SLAs, and error budgets.
• Strong automation capabilities aimed at minimizing toil through self-healing systems and infrastructure automation.
• Advanced skills in Python, Bash, and/or PowerShell for tooling and automation purposes.
• Profound understanding of networking concepts, including TCP/IP, DNS, load balancing, VPNs, and cloud-native networking.
• Background in Site Reliability Engineering, DevOps, or Infrastructure roles, with experience leading engineering teams.
• Demonstrated success in leading incident response efforts and implementing reliability enhancements.
• Annual incentive bonus.
• Country-specific benefits.
• Support for disability and accommodations during the hiring process.
Koniag Government Services
FP Markets (First Prudential Markets)
Modern Campus
InRule
Get handpicked remote jobs straight to your inbox weekly.