
Staff Field Reliability Engineer
Posted Aug 12

Posted Aug 12
This is a fully remote position, open to applicants in United States.
• Establish architectural frameworks and operational standards for Refinery as a Service and Honeycomb Private Cloud across various AWS accounts and regions.
• Design Terraform modules, Helm charts, and automate deployment processes.
• Set the technical course for instrumentation, monitoring, and management of infrastructure.
• Take ownership of capacity planning, scaling strategies, upgrade sequencing, and cost optimization in multi-region AWS environments.
• Develop platforms and automation that empower the FRE team to scale effectively.
• Act as the final technical escalation point for unique, high-stakes customer scenarios.
• Address infrastructure and observability challenges involving distributed systems, Kubernetes, AWS networking, and polyglot service meshes.
• Collaborate with customer SRE, platform, and engineering leadership on escalations and architectural redesigns.
• Provide senior incident management for managed services.
• Create playbooks, diagnostic frameworks, and tools.
• Influence Honeycomb's OpenTelemetry open-source strategy and represent Honeycomb in OpenTelemetry SIGs.
• Develop reference architectures and integration documentation.
• Lead contributions to Honeycomb's open-source initiatives.
• Serve as the final technical authority for Solutions Architects on complex deals and production troubleshooting.
• Facilitate architecture reviews, SLO workshops, instrumentation deep-dives, POCs, and pilot programs.
• Propel roadmap prioritization based on strategic customer requirements.
• Construct internal tools and user interfaces for the FRE function.
• Foster alignment among Solutions Architecture, Customer Success, Support, Product, and Engineering.
• Manage trade-offs across the FRE charter.
• Mentor IC3/IC4 engineers regarding technical scope and career advancement.
• Engage with Honeycomb leadership and customer C-suite.
• 9+ years of experience in engineering, SRE, infrastructure, DevOps, or a similar field, demonstrating Staff-level technical scope and impact.
• Extensive hands-on experience with Kubernetes, preferably with EKS.
• Strong expertise in AWS services including EC2, EKS, ECS, ALB/NLB, VPC, PrivateLink, IAM, S3, and Route53.
• Proficient in multi-account architecture design, service quotas, and strategies for cost optimization.
• Experience in senior incident management, including incident response, triage, and postmortem process enhancements.
• Mastery of Infrastructure as Code practices using Terraform, Helm, Chef, and Ansible.
• In-depth knowledge of observability principles, including structured logging, distributed tracing, metrics, SLOs/SLIs, and the instrumentation lifecycle.
• Strong understanding of OpenTelemetry SDK, Collector architecture, processors, exporters, and semantic conventions, with active public community contributions and leadership.
• Proficiency in at least two programming languages such as Go, Python, Java, TypeScript/Node.js, and .NET.
• Exceptional executive communication abilities.
• Capability to set direction in ambiguous, high-pressure environments and create structures that enable independent operation.
• Visa sponsorship or transfers are not currently available.
• Must confirm identity and eligibility to work.
• Generous equity through an employee-friendly stock program.
• Transparent compensation based on experience levels.
• Unlimited paid time off (PTO).
• A distributed-first culture and mindset.
• Stipends for home office, co-working, and internet expenses.
• Comprehensive benefits coverage for employees, with additional options for dependents.
• Up to 16 weeks of paid parental leave, regardless of the path to parenthood.
• Annual development allowance.
DATAGROUP
Ambush
DuoKey
TEKsystems
Get handpicked remote jobs straight to your inbox weekly.