Lead Site Reliability Engineer

Posted Sep 2

This is a fully remote position, open to applicants in United States.

📋 Description

• Lead and manage initiatives for infrastructure modernization, transitioning from legacy compute environments to container-orchestrated infrastructure.

• Architect and uphold infrastructure as code across various cloud platforms.

• Enhance CI/CD systems and deployment tools to ensure efficient, observable, and recoverable releases.

• Provide technical leadership through mentorship, architectural guidance, knowledge sharing, and support for engineering practices.

• Establish and refine SLI and SLO practices, along with monitoring, alerting, and load-testing capabilities.

• Oversee platform incident response, troubleshooting, root cause analysis, and implementation of corrective measures.

• Increase transparency around cloud infrastructure costs and integrate cost considerations into architectural choices.

• Elevate the developer experience by improving infrastructure, development environments, deployment workflows, and production feedback loops.

• Collaborate with Security and Engineering teams on access controls, infrastructure hardening, compliance, and secure infrastructure practices.

• Identify operational and infrastructure risks, suggest priorities, and assist in shaping the Systems Engineering technical roadmap.

• Mentor engineers and contribute to the growth of the Systems Engineering team and its technical practices.

• Carry out additional duties as assigned.


⛳️ Requirements

• A minimum of 8 years of hands-on experience in infrastructure, DevOps, platform engineering, site reliability engineering, or a related field, including experience in building and operating production systems.

• Extensive hands-on experience in designing, operating, and troubleshooting highly available production infrastructure.

• Production experience across multiple major cloud providers, with substantial expertise in at least one of AWS or Google Cloud and proficiency in the other.

• Significant experience with container orchestration and Kubernetes in production settings, including cluster operations, workload configuration, reliability, and troubleshooting.

• Proven experience in modernizing production infrastructure, including migrations from VM-based or legacy environments to containerized or cloud-native architectures.

• Strong infrastructure-as-code expertise using Terraform or similar tools, focusing on repeatability and automation.

• Experience in building, operating, or significantly enhancing CI/CD systems and deployment infrastructure.

• Excellent incident response and troubleshooting skills, with experience diagnosing complex distributed-system failures and contributing to effective post-incident reviews.

• Proficient Linux administration skills and programming or scripting capabilities in Ruby, Python, or a comparable language.

• Experience in a SaaS environment where reliability, availability, and production stability are paramount.

• Ability to effectively collaborate with distributed teams across US and European time zones and participate in an on-call rotation.

• Strong communication abilities, with the capacity to provide technical direction, mentor other engineers, and influence infrastructure decisions across teams.

• A Bachelor's degree in a relevant field or equivalent practical experience; equivalent experience is genuinely accepted for this role.

• Experience in production environments across both AWS and Google Cloud simultaneously is preferred.

• Prior experience in leadership or management, cloud cost management or FinOps, Spinnaker/Jenkins, Ruby on Rails, SOC 2, AI-assisted development tools, or learning technology is preferred.


🏝️ Benefits

• Medical - 100% of employee premiums for selected individual plans.

• Dental - 100% of employee premiums covered.

• Vision - 100% of employee premiums covered.

• Access to LinkedIn Learning.

• 401(k) plan with matching (US Based Only).

• Flexible PTO.

• Subscription to Calm.

• Annual Company Retreat.

• Personal development budgets.

People also viewed

General Dynamics Information Technology7 hours ago

Principal DevSecOps Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$144.5k – $195.5k/year
ApplyView job
VALCE Talent Solutions7 hours ago

DevOps/SRE Engineer

MX flagMexico OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
VALCE Talent Solutions10 hours ago

Senior Release Engineer

MX flagMexico OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
ImmunityBio, Inc.10 hours ago

DevOps Engineer

US flagCalifornia OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$130.5k – $150k/year
ApplyView job
OXIO11 hours ago

Site Reliability Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Rimutee13 hours ago

DevOps Engineer – LATAM

Latin AmericaFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers