
Staff Software Engineer – L4
Posted 5 days ago

Posted 5 days ago
This is a fully remote position, open to applicants in Ireland.
• Take charge of the reliability standards for production services, encompassing availability, latency, capacity, efficiency, performance, monitoring, and alerting.
• Establish, implement, and operate with SLIs and SLOs, leveraging error budgets to shape engineering priorities.
• Recognize stability risks and proactively mitigate them to prevent customer impact.
• Minimize repair items and avert the recurrence of incident classes.
• Enhance detection, response, and recovery times.
• Fortify failure domains, validate recovery paths, and ensure safer production changes.
• Engage in on-call duties and lead responses during production degradation.
• Compose post-mortems, ascertain root causes, and drive subsequent actions to completion.
• Identify, troubleshoot, report, and document production issues.
• Develop, configure, and deploy maintainable, reviewed, documented, and tested code that boosts reliability.
• Coordinate complex changes across systems and services, including design modifications, technical decisions, migrations, and upgrades.
• Lead debugging, troubleshooting, and assessment of service architecture and design.
• Utilize code reviews to enhance code quality.
• Decrease operational overhead for infrastructure and services.
• Propel cross-team projects from initiation to completion.
• Oversee programs, estimate and convey delivery timelines, track milestones, and keep stakeholders informed.
• Identify and communicate any changes that may impact stability.
• A minimum of 8 years of relevant engineering experience, with a significant focus on reliability, infrastructure, or platform engineering.
• Proven responsibility for production systems, including having managed a pager for critical services and owning outcomes during failures.
• Strong foundation in software engineering and the capability to build and deploy production code.
• Experience in defining and operating SLIs and SLOs, utilizing error budgets to guide engineering priorities.
• Extensive background in production operations, including incident command, post-mortem analysis, capacity planning, and observability.
• A history of preventing recurrence by minimizing incident classes and operational toil.
• Experience in driving changes across multiple teams and achieving alignment without formal authority.
• Proven track record of enhancing engineers through code review, design feedback, and mentorship.
• Familiarity with large-scale distributed systems in a cloud setting.
• Knowledge of infrastructure-as-code, container orchestration, and GitOps-style delivery.
• Experience with multi-region architecture, failure-domain design, or regional expansion initiatives.
• Background in chaos engineering, game days, or other proactive resilience validation methods.
• Competitive salary.
• Generous vacation policy.
• Extensive parental and wellness leave.
• Comprehensive healthcare coverage.
• Retirement savings plan.
• Support for employee volunteering and donation initiatives.
Malbek
ElevenLabs
ElevenLabs
Solidus Labs
Get handpicked remote jobs straight to your inbox weekly.