Remotery

Staff Software Engineer – Reliability & Platform

Posted 5 hours ago

This is a fully remote position, open to applicants in United Kingdom.

📋 Description

• Take charge of production operability by troubleshooting complex issues, enhancing system visibility, and addressing recurring problems at their source.

• Manage production health for services — overseeing everything from detection to resolution and prevention.

• Enhance mean time to detect (MTTD), mean time to resolve (MTTR), and reduce recurrence rates for issues.

• Identify systemic challenges and eliminate recurring issues through code improvements, architectural enhancements, and more effective operational tools.

• Boost observability across services — including logs, metrics, and alerting — to facilitate quicker diagnosis and resolution.

• Design and refine debugging workflows, runbooks, and internal tools for engineers.

• Alleviate operational burdens by simplifying systems for better understanding, operation, and troubleshooting.

• Collaborate closely with product teams to integrate production insights back into design and development processes.

• Mitigate support and incident workloads by addressing root causes and enhancing system design, rather than merely resolving isolated issues.


⛳️ Requirements

• Over 4 years of software engineering experience with ownership of production systems, reliability, or operational enhancements.

• Strong backend development expertise (Ruby, Node.js, or a similar language).

• Comprehensive understanding of REST APIs, service contracts, and software design principles.

• Experience working across backend services, APIs, and production systems.

• Familiarity with building and managing services in AWS or comparable cloud environments.

• Experience with observability, production debugging, and incident response.

• Willingness to engage in a mandatory 24/7 on-call rotation.

• Background in responding to production incidents and contributing to reliability enhancements.

• Experience with, or a strong interest in, AI-powered development tools (e.g., GitHub Copilot, ChatGPT, Cursor).


🏝️ Benefits

• Competitive salary

• Flexible working hours

• Professional development budget

• Home office setup allowance

• Global team events

People also viewed

Fundraise Up4 hours ago

Senior DevOps Engineer

PL flagPoland OnlyFull-timeDevOps & Site Reliability Engineer (SRE)€6,000 – €6,800/month
ApplyView job
Empower4 hours ago

Lead Site Reliability Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$114k – $165.3k/year
ApplyView job
Harrods4 hours ago

DevOps Manager

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Aufinity Group | España4 hours ago

Software Developer – DevOps

DE flagGermany OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Zipdev4 hours ago

Senior Site Reliability Engineer

BR flagBrazil OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Valtech5 hours ago

Senior Site Reliability Engineer

MK flagMacedonia OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers