
Lead Site Reliability Engineer
Posted 3 hours ago

Posted 3 hours ago
This is a fully remote position, open to applicants in United States.
• Drive cross-functional reliability projects across various value streams and oversee their execution across multiple teams.
• Define and enhance SRE best practices, tools, and methodologies throughout the organization.
• Design and implement enterprise-scale, multi-region AWS infrastructure that optimizes reliability, cost, performance, and security.
• Establish and manage SLOs, SLIs, and error budgets for essential services, utilizing them to influence prioritization decisions.
• Act as incident commander during major incidents and lead postmortems that yield actionable items and organizational insights.
• Oversee disaster recovery planning for critical financial services infrastructure.
• Develop shared Infrastructure as Code foundations using Terraform, including reusable modules, standards, and patterns embraced by teams.
• Create and implement production-scale Kubernetes patterns, covering multi-tenancy, security policies, and advanced scheduling techniques.
• Set observability standards and strategies using Datadog and Splunk (metrics, logging, tracing, dashboards, and alerting).
• Establish CI/CD standards and methodologies, including pipeline-as-code and large-scale progressive delivery.
• Lead chaos engineering, game days, and structured reliability testing initiatives.
• Propel FinOps initiatives to optimize cloud expenditures while achieving reliability objectives.
• Guide a functional team of SREs (without direct reports) on various projects and operational activities.
• Provide mentorship to SREs across multiple levels through coaching, design reviews, code reviews, and training sessions.
• Collaborate with Engineering, Product, and Security leaders to align reliability efforts with business priorities, zero-trust architecture, and compliance requirements.
• Bachelor’s degree in Computer Science, Information Technology, or a related field (or equivalent practical experience).
• 7 to 10 years of Site Reliability Engineering experience (or similar), demonstrating technical leadership.
• Proven capability to lead technical teams and manage complex projects to successful completion.
• In-depth AWS expertise, including the design of large-scale, multi-region architectures.
• Advanced Kubernetes knowledge, encompassing features, security, and operations at a production scale.
• Expertise in Infrastructure as Code utilizing Terraform, including the creation of shared platforms and frameworks.
• Strong software engineering proficiency with production experience in Python and/or Go.
• Extensive experience with observability platforms (Datadog, Splunk) and scaling monitoring implementations.
• Comprehensive understanding of CI/CD principles and experience in deploying enterprise-grade pipelines.
• Established history of leading major incidents and conducting thorough postmortems.
• Solid understanding of security, networking, and infrastructure design patterns.
• Excellent communication skills with the ability to convey complex technical concepts to varied audiences.
• Experience in mentoring engineers and fostering technical growth within teams.
• Medical, dental, vision, and life insurance.
• Retirement savings – 401(k) plan featuring generous company matching contributions (up to 6%), financial advisory services, potential discretionary company contributions, and a wide range of investment options.
• Tuition reimbursement up to $5,250 per year.
• Business-casual work environment with the option to wear jeans.
• Generous paid time off upon hire, which includes a paid time off program, ten paid company holidays, and three floating holidays each calendar year.
• Paid volunteer time – 16 hours per calendar year.
• Leave of absence programs, including paid parental leave, paid short- and long-term disability, and Family and Medical Leave (FMLA).
• Business Resource Groups (BRGs) – BRGs promote inclusion and collaboration across our organization and within the communities we serve. They are open to all employees.
Fundraise Up
Harrods
Aufinity Group | España
Zipdev
Get handpicked remote jobs straight to your inbox weekly.