
Product Reliability Engineer
Posted Jul 28

Posted Jul 28
This is a fully remote position, open to applicants in Czechia.
• Collaborate directly with customers as well as our Solution Architecture and Customer Success teams on L2/L3 escalations—sharing insights, conducting root-cause analysis, and resolving intricate issues related to packaging, deployment, upgrades, and runtime across diverse Kubernetes environments.
• Propel issues toward resolution by replicating problems locally, pinpointing root causes, and coordinating solutions with engineering—subsequently documenting insights in concise RCAs that translate into actionable enhancements.
• Develop and sustain diagnostic tools including support bundles, health checks, environment validators, and "what changed?" aids that expedite future troubleshooting significantly.
• Take ownership of the test automation infrastructure roadmap, enhancing CI stability, minimizing flaky tests, and establishing reproducible integration/e2e environments that identify issues before they reach customers.
• Set up and uphold performance baselines and regression tests that act as actionable checkpoints, assisting teams in detecting scale and latency problems early on.
• Enhance installation and upgrade reliability by identifying persistent failure modes and eradicating them through product modifications, automation, and safeguards.
• Produce production-quality code in Python, Go, or Rust for internal tools and product enhancements that directly improve reliability.
• Complete the reliability feedback loop by systematically converting field issues into improved tests, observability, documentation, and product defaults—measuring success through reduced time-to-resolution and fewer recurring incidents.
• 4-7 years of experience in production engineering, SRE, platform engineering, or similar roles where you have managed reliability and customer escalations.
• Strong foundation in software engineering principles including design, debugging, testing, code review, and a focus on maintainable, production-quality code.
• Practical expertise in Kubernetes sufficient to debug actual deployments: troubleshooting resources, networking, storage, RBAC, and platform-specific nuances across various distributions.
• Profound troubleshooting instincts and observability experience utilizing logs, metrics, and traces to swiftly diagnose issues in complex, distributed systems.
• Experience with at least one of the following: Python, Go, or Rust for developing tools and contributing to product code (expertise in all three is not required).
• Exceptional problem decomposition and communication skills—you can untangle messy, ambiguous issues and clearly articulate your findings and recommendations.
• Self-motivated remote work capability with strong asynchronous communication skills and the ability to work independently in a fast-paced environment where priorities shift based on customer demands.
• Collaborative mindset with experience working across product, engineering, and customer-facing teams to drive systematic enhancements.
• The people: Collaborate alongside top-tier engineers who have built and scaled automation platforms in production. Encounter daily technical challenges with intelligent colleagues who encourage your growth.
• The product: Influence Infrahub based on genuine customer needs. Your contributions directly affect features, integrations, and roadmap priorities.
• The mission: We are making enterprise-grade infrastructure automation accessible to all organizations. Open-source at the core, production-ready out of the box. This is a multi-year journey, not just a quarterly sprint.
• The impact: You'll engage with teams managing some of the world's most intricate infrastructure deployments, addressing challenges that resonate throughout entire organizations.
• Our Commitment to Diversity and Inclusion: OpsMill is dedicated to fostering a diverse and inclusive team. We believe that varied perspectives strengthen our innovation. We welcome applications from candidates of all backgrounds and experiences and are committed to creating an inclusive environment where everyone can excel.
DATAGROUP
Ambush
DuoKey
TEKsystems
Get handpicked remote jobs straight to your inbox weekly.