
Product Reliability Engineer
Posted Jul 28

Posted Jul 28
This is a fully remote position, open to applicants in United States.
• Collaborate closely with customers as well as our Solution Architecture and Customer Success teams on L2/L3 escalations—articulating findings, facilitating root-cause analysis, and addressing intricate packaging, deployment, upgrade, and runtime challenges across diverse Kubernetes environments.
• Lead the resolution of issues by replicating problems locally, pinpointing root causes, and coordinating solutions with engineering—while documenting insights in clear RCAs that lead to actionable enhancements.
• Develop and sustain diagnostic tools such as support bundles, health checks, environment validators, and "what changed?" utilities that significantly accelerate future troubleshooting.
• Take ownership of the test automation infrastructure roadmap, enhancing CI stability, minimizing flaky tests, and establishing reproducible integration/e2e environments that identify issues before they affect customers.
• Set up and uphold performance baselines and regression tests that act as actionable checkpoints, enabling teams to detect scale and latency concerns early.
• Enhance installation and upgrade reliability by recognizing recurring failure patterns and eliminating them via product modifications, automation, and safeguards.
• Write high-quality production code in Python, Go, or Rust for internal tools and product enhancements that directly improve reliability.
• Complete the reliability feedback loop by systematically transforming field issues into improved tests, observability, documentation, and product defaults—tracking success through shorter time-to-resolution and reduced repeat incidents.
• 4-7 years of experience in production engineering, SRE, platform engineering, or comparable roles where you have taken ownership of reliability and customer escalations.
• Strong foundational software engineering skills including design, debugging, testing, code review, and a commitment to maintainable, production-quality code.
• Practical expertise in Kubernetes sufficient to debug actual deployments: troubleshooting resources, networking, storage, RBAC, and platform-specific nuances across various distributions.
• Strong troubleshooting instincts and observability experience utilizing logs, metrics, and traces to swiftly diagnose issues in complex, distributed systems.
• Proficiency in at least one of the following: Python, Go, or Rust for developing tools and contributing to product code (expertise in all three is not necessary).
• Exceptional problem decomposition and communication abilities—you can clarify messy, ambiguous issues and articulate your findings and recommendations effectively.
• Ability to work independently in a remote setting with strong asynchronous communication skills, thriving in a fast-paced environment where priorities shift based on customer needs.
• A collaborative spirit with experience working across product, engineering, and customer-facing teams to drive systematic enhancements.
• The people: Collaborate with world-class engineers who have built and scaled automation platforms in production. Engage in daily technical challenges with intelligent colleagues who inspire you to grow.
• The product: Influence Infrahub based on genuine customer needs. Your contributions directly impact features, integrations, and roadmap priorities.
• The mission: We aim to make enterprise-grade infrastructure automation accessible to every organization. Open-source at its core, production-ready out of the box. This is a long-term journey, not a quarterly sprint.
• The impact: You'll collaborate with teams managing some of the world's most intricate infrastructure deployments, addressing challenges that resonate throughout entire organizations.
• Our Commitment to Diversity and Inclusion: OpsMill is dedicated to fostering a diverse and inclusive team. We believe that varied perspectives strengthen us and foster innovation. We welcome applications from candidates of all backgrounds and experiences, and we are committed to providing an inclusive environment where everyone can excel.
DATAGROUP
Ambush
DuoKey
TEKsystems
Get handpicked remote jobs straight to your inbox weekly.