
Staff Software Engineer – Reliability, Platform
Posted Jul 22

Posted Jul 22
This is a fully remote position, open to applicants in Hungary.
• Take ownership of production operability by troubleshooting complex issues, enhancing system visibility, and addressing recurring problems at their origin.
• Enhance mean time to detect (MTTD), mean time to resolve (MTTR), and recurrence rates of issues.
• Detect systemic issues and eliminate recurring problems through code modifications, architecture enhancements, and improved operational tools.
• Advance observability across services — including logs, metrics, and alerting — for quicker diagnosis and resolution.
• Design and refine debugging workflows, runbooks, and internal tools for engineers.
• Lessen operational strain by making systems simpler to comprehend, operate, and troubleshoot.
• Collaborate closely with product teams to incorporate production insights back into design and development.
• Mitigate support and incident workload by addressing root causes and enhancing system design rather than merely fixing individual issues.
• Over 4 years of software engineering experience with responsibility for production systems, reliability, or operational enhancements.
• Proficient backend development experience (Ruby, Node.js, or similar technologies).
• Strong grasp of REST APIs, service contracts, and software design principles.
• Experience in building and managing services within AWS or comparable cloud environments.
• Good understanding of distributed systems, cloud-native architecture, and CI/CD processes.
• Background in observability, production debugging, and incident management.
• Willingness to engage in a mandatory 24/7 on-call rotation.
• Experience with, or a keen interest in, AI-driven development tools (e.g., GitHub Copilot, ChatGPT, Cursor).
• Health insurance
• Professional development opportunities
Quantiphi
Encompass Corporation
Shippit
Group O
Get handpicked remote jobs straight to your inbox weekly.