Remotery

Lead Site Reliability Engineer

Posted 3 hours ago

This is a fully remote position, open to applicants in United States.

📋 Description

• Drive cross-functional reliability projects across various value streams and oversee their execution across multiple teams.

• Define and enhance SRE best practices, tools, and methodologies throughout the organization.

• Design and implement enterprise-scale, multi-region AWS infrastructure that optimizes reliability, cost, performance, and security.

• Establish and manage SLOs, SLIs, and error budgets for essential services, utilizing them to influence prioritization decisions.

• Act as incident commander during major incidents and lead postmortems that yield actionable items and organizational insights.

• Oversee disaster recovery planning for critical financial services infrastructure.

• Develop shared Infrastructure as Code foundations using Terraform, including reusable modules, standards, and patterns embraced by teams.

• Create and implement production-scale Kubernetes patterns, covering multi-tenancy, security policies, and advanced scheduling techniques.

• Set observability standards and strategies using Datadog and Splunk (metrics, logging, tracing, dashboards, and alerting).

• Establish CI/CD standards and methodologies, including pipeline-as-code and large-scale progressive delivery.

• Lead chaos engineering, game days, and structured reliability testing initiatives.

• Propel FinOps initiatives to optimize cloud expenditures while achieving reliability objectives.

• Guide a functional team of SREs (without direct reports) on various projects and operational activities.

• Provide mentorship to SREs across multiple levels through coaching, design reviews, code reviews, and training sessions.

• Collaborate with Engineering, Product, and Security leaders to align reliability efforts with business priorities, zero-trust architecture, and compliance requirements.


⛳️ Requirements

• Bachelor’s degree in Computer Science, Information Technology, or a related field (or equivalent practical experience).

• 7 to 10 years of Site Reliability Engineering experience (or similar), demonstrating technical leadership.

• Proven capability to lead technical teams and manage complex projects to successful completion.

• In-depth AWS expertise, including the design of large-scale, multi-region architectures.

• Advanced Kubernetes knowledge, encompassing features, security, and operations at a production scale.

• Expertise in Infrastructure as Code utilizing Terraform, including the creation of shared platforms and frameworks.

• Strong software engineering proficiency with production experience in Python and/or Go.

• Extensive experience with observability platforms (Datadog, Splunk) and scaling monitoring implementations.

• Comprehensive understanding of CI/CD principles and experience in deploying enterprise-grade pipelines.

• Established history of leading major incidents and conducting thorough postmortems.

• Solid understanding of security, networking, and infrastructure design patterns.

• Excellent communication skills with the ability to convey complex technical concepts to varied audiences.

• Experience in mentoring engineers and fostering technical growth within teams.


🏝️ Benefits

• Medical, dental, vision, and life insurance.

• Retirement savings – 401(k) plan featuring generous company matching contributions (up to 6%), financial advisory services, potential discretionary company contributions, and a wide range of investment options.

• Tuition reimbursement up to $5,250 per year.

• Business-casual work environment with the option to wear jeans.

• Generous paid time off upon hire, which includes a paid time off program, ten paid company holidays, and three floating holidays each calendar year.

• Paid volunteer time – 16 hours per calendar year.

• Leave of absence programs, including paid parental leave, paid short- and long-term disability, and Family and Medical Leave (FMLA).

• Business Resource Groups (BRGs) – BRGs promote inclusion and collaboration across our organization and within the communities we serve. They are open to all employees.

People also viewed

Fundraise Up3 hours ago

Senior DevOps Engineer

PL flagPoland OnlyFull-timeDevOps & Site Reliability Engineer (SRE)€6,000 – €6,800/month
ApplyView job
Harrods3 hours ago

DevOps Manager

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Aufinity Group | España3 hours ago

Software Developer – DevOps

DE flagGermany OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Zipdev3 hours ago

Senior Site Reliability Engineer

BR flagBrazil OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Valtech4 hours ago

Senior Site Reliability Engineer

MK flagMacedonia OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
One Identity4 hours ago

Staff Software Engineer – Reliability & Platform

GB flagUnited Kingdom OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers