Remotery

Staff Field Reliability Engineer

Posted Aug 12

This is a fully remote position, open to applicants in United States.

📋 Description

• Establish architectural frameworks and operational standards for Refinery as a Service and Honeycomb Private Cloud across various AWS accounts and regions.

• Design Terraform modules, Helm charts, and automate deployment processes.

• Set the technical course for instrumentation, monitoring, and management of infrastructure.

• Take ownership of capacity planning, scaling strategies, upgrade sequencing, and cost optimization in multi-region AWS environments.

• Develop platforms and automation that empower the FRE team to scale effectively.

• Act as the final technical escalation point for unique, high-stakes customer scenarios.

• Address infrastructure and observability challenges involving distributed systems, Kubernetes, AWS networking, and polyglot service meshes.

• Collaborate with customer SRE, platform, and engineering leadership on escalations and architectural redesigns.

• Provide senior incident management for managed services.

• Create playbooks, diagnostic frameworks, and tools.

• Influence Honeycomb's OpenTelemetry open-source strategy and represent Honeycomb in OpenTelemetry SIGs.

• Develop reference architectures and integration documentation.

• Lead contributions to Honeycomb's open-source initiatives.

• Serve as the final technical authority for Solutions Architects on complex deals and production troubleshooting.

• Facilitate architecture reviews, SLO workshops, instrumentation deep-dives, POCs, and pilot programs.

• Propel roadmap prioritization based on strategic customer requirements.

• Construct internal tools and user interfaces for the FRE function.

• Foster alignment among Solutions Architecture, Customer Success, Support, Product, and Engineering.

• Manage trade-offs across the FRE charter.

• Mentor IC3/IC4 engineers regarding technical scope and career advancement.

• Engage with Honeycomb leadership and customer C-suite.


⛳️ Requirements

• 9+ years of experience in engineering, SRE, infrastructure, DevOps, or a similar field, demonstrating Staff-level technical scope and impact.

• Extensive hands-on experience with Kubernetes, preferably with EKS.

• Strong expertise in AWS services including EC2, EKS, ECS, ALB/NLB, VPC, PrivateLink, IAM, S3, and Route53.

• Proficient in multi-account architecture design, service quotas, and strategies for cost optimization.

• Experience in senior incident management, including incident response, triage, and postmortem process enhancements.

• Mastery of Infrastructure as Code practices using Terraform, Helm, Chef, and Ansible.

• In-depth knowledge of observability principles, including structured logging, distributed tracing, metrics, SLOs/SLIs, and the instrumentation lifecycle.

• Strong understanding of OpenTelemetry SDK, Collector architecture, processors, exporters, and semantic conventions, with active public community contributions and leadership.

• Proficiency in at least two programming languages such as Go, Python, Java, TypeScript/Node.js, and .NET.

• Exceptional executive communication abilities.

• Capability to set direction in ambiguous, high-pressure environments and create structures that enable independent operation.

• Visa sponsorship or transfers are not currently available.

• Must confirm identity and eligibility to work.


🏝️ Benefits

• Generous equity through an employee-friendly stock program.

• Transparent compensation based on experience levels.

• Unlimited paid time off (PTO).

• A distributed-first culture and mindset.

• Stipends for home office, co-working, and internet expenses.

• Comprehensive benefits coverage for employees, with additional options for dependents.

• Up to 16 weeks of paid parental leave, regardless of the path to parenthood.

• Annual development allowance.

People also viewed

DATAGROUP2 days ago

DevOps Engineer

DE flagGermany OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Ambush2 days ago

DevOps Engineer

BR flagBrazil OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
DuoKey2 days ago

DevOps Engineer

MU flagMauritius OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
TEKsystems3 days ago

SRE – CloudOps, Practice Architect II

US flagIllinois OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$148.2k – $222.4k/year
ApplyView job
TEKsystems3 days ago

SRE CloudOps Practice Architect II

US flagTexas OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$148.2k – $222.4k/year
ApplyView job
Level Data3 days ago

Senior DevOps Engineer

US flagMassachusetts OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$120k – $135k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers