Senior Site Reliability Engineer

Posted 8 hours ago

This is a fully remote position, open to applicants in United States, +2 more locations.

📋 Description

• Enhance production reliability and bolster system resilience within a Site Reliability Engineering (SRE) focused team.

• Advocate for high standards and adherence to industry best practices.

• Facilitate communication with teams and stakeholders throughout various projects.

• Introduce innovative ideas and inspire others to do the same.

• Analyze and resolve intricate technical issues.

• Collaborate across diverse technologies in a rapidly evolving industry.

• Take part in on-call rotations, incident responses, and conduct blameless post-incident reviews.

• Write code, manage alerts, enhance solutions, and provide support to colleagues.

• Involve stakeholders in requirements analysis and demonstrations.

• Ensure systems are secure, maintainable, and consistently available.

• Contribute to customer success and align with company objectives.


⛳️ Requirements

• Over 5 years of experience managing Linux systems and associated infrastructure in production settings.

• Familiarity with SLIs, SLOs, SLAs, error budgets, blast radius, and conducting blameless postmortems.

• Focused on automation, minimizing toil, and averting problem recurrence.

• Proven history of writing runbooks for wider teams.

• Strong foundational knowledge in Kubernetes and its broader ecosystem.

• Experience with cloud infrastructure; AWS is highly preferred.

• Bare-metal experience is an additional advantage.

• Proficient in tool development using Bash, with Python or Go preferred, or similar languages.

• Experience with infrastructure-as-code tools; Terraform is preferred.

• Familiarity with CI/CD and version control systems; GitHub is preferred.

• Experience with databases such as Postgres, Cassandra, or ClickHouse is a plus.

• Experience in operating production observability stacks that encompass metrics, logs, and traces.

• Strong troubleshooting skills and a proactive approach to incident response.

• A history of ongoing professional development.

• A self-directed work style that suits an asynchronous, globally distributed team.

• Willingness to take on adjacent tasks when necessary.


🏝️ Benefits

• Flexible working environment – embracing a remote-first culture with coworking options available.

• Four weeks of paid annual leave.

• Parental leave.

• Birthday leave.

• Option to purchase additional annual leave.

• Wellness allowance.

• Initiatives focused on employee wellbeing.

• Generous study and training budget.

• Five days of paid study leave.

• Creative and modern workspaces.

• Recognition programs, including Legend and Kudos awards.

People also viewed

Capstone Integrated Solutions6 hours ago

AWS DevOps Engineer, MLOps

US flagNew York OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Centric Software6 hours ago

Site Reliability Engineer

PT flagPortugal OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Independence Pet Group7 hours ago

DevOps Engineer

US flagNew York OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Zocdoc7 hours ago

Senior Site Reliability Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$180k – $220k/year
ApplyView job
Fairsource7 hours ago

DevOps, Kubernetes Consultant

DE flagGermany OnlyFull-timeDevOps & Site Reliability Engineer (SRE)€110k – €140k/year
ApplyView job
GFT Technologies9 hours ago

Ingeniero DevOps Sr.

CR flagCosta Rica OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers