Staff Software Engineer – L4

Posted 5 days ago

This is a fully remote position, open to applicants in Ireland.

📋 Description

• Take charge of the reliability standards for production services, encompassing availability, latency, capacity, efficiency, performance, monitoring, and alerting.

• Establish, implement, and operate with SLIs and SLOs, leveraging error budgets to shape engineering priorities.

• Recognize stability risks and proactively mitigate them to prevent customer impact.

• Minimize repair items and avert the recurrence of incident classes.

• Enhance detection, response, and recovery times.

• Fortify failure domains, validate recovery paths, and ensure safer production changes.

• Engage in on-call duties and lead responses during production degradation.

• Compose post-mortems, ascertain root causes, and drive subsequent actions to completion.

• Identify, troubleshoot, report, and document production issues.

• Develop, configure, and deploy maintainable, reviewed, documented, and tested code that boosts reliability.

• Coordinate complex changes across systems and services, including design modifications, technical decisions, migrations, and upgrades.

• Lead debugging, troubleshooting, and assessment of service architecture and design.

• Utilize code reviews to enhance code quality.

• Decrease operational overhead for infrastructure and services.

• Propel cross-team projects from initiation to completion.

• Oversee programs, estimate and convey delivery timelines, track milestones, and keep stakeholders informed.

• Identify and communicate any changes that may impact stability.


⛳️ Requirements

• A minimum of 8 years of relevant engineering experience, with a significant focus on reliability, infrastructure, or platform engineering.

• Proven responsibility for production systems, including having managed a pager for critical services and owning outcomes during failures.

• Strong foundation in software engineering and the capability to build and deploy production code.

• Experience in defining and operating SLIs and SLOs, utilizing error budgets to guide engineering priorities.

• Extensive background in production operations, including incident command, post-mortem analysis, capacity planning, and observability.

• A history of preventing recurrence by minimizing incident classes and operational toil.

• Experience in driving changes across multiple teams and achieving alignment without formal authority.

• Proven track record of enhancing engineers through code review, design feedback, and mentorship.

• Familiarity with large-scale distributed systems in a cloud setting.

• Knowledge of infrastructure-as-code, container orchestration, and GitOps-style delivery.

• Experience with multi-region architecture, failure-domain design, or regional expansion initiatives.

• Background in chaos engineering, game days, or other proactive resilience validation methods.


🏝️ Benefits

• Competitive salary.

• Generous vacation policy.

• Extensive parental and wellness leave.

• Comprehensive healthcare coverage.

• Retirement savings plan.

• Support for employee volunteering and donation initiatives.

People also viewed

Malbek18 hours ago

Full Stack Developer

US flagUnited States OnlyFull-timeFull-stack Engineer
ApplyView job
ElevenLabs20 hours ago

Forward Deployed Engineer – Software Engineer

MX flagMexico OnlyFull-timeFull-stack Engineer
ApplyView job
ElevenLabs22 hours ago

Forward Deployed Engineer – Software Engineer

GB flagUnited Kingdom OnlyFull-timeFull-stack Engineer
ApplyView job
Solidus Labs22 hours ago

Tech Lead, Transaction Monitoring

GB flagUnited Kingdom OnlyFull-timeFull-stack Engineer
ApplyView job
Marigold1 day ago

Full Stack Software Engineer

AU flagAustralia OnlyFull-timeFull-stack Engineer
ApplyView job
NationsBenefits1 day ago

Software Development Engineer II

US flagFlorida OnlyFull-timeFull-stack Engineer
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers