
Lead Observability Engineer
Posted 1 day ago

Posted 1 day ago
This is a fully remote position, open to applicants in United States.
β’ Serve as the primary contact for customer engineering leadership while leading an embedded TechPod as Pod Leader.
β’ Conduct a two-month Observability Maturity Assessment.
β’ Evaluate the observability landscape across various vendors, agents, collectors, query surfaces, data volumes, and operating models.
β’ Confirm metric cardinality, active series, scrape target health, log volumes, trace sampling, retention, and telemetry fidelity.
β’ Develop total cost of ownership comparisons for current and future states.
β’ Design an AWS-native observability target architecture utilizing OpenTelemetry and AWS services.
β’ Collaborate with Security, Legal, and Engineering teams on telemetry classification and policy-driven retention.
β’ Map existing capabilities to AWS-native equivalents and pinpoint gaps.
β’ Define and execute performance-parity tests for query performance, alert latency, and data fidelity.
β’ Plan and oversee migration execution, which includes pipeline cutover, porting log transformations to OpenTelemetry, rebuilding dashboards, and deduplicating/rebuilding alerts.
β’ Establish OpenTelemetry conventions, collector topology, and sampling strategies.
β’ Revise service ownership tagging for telemetry cost attribution.
β’ Analyze platform expenditures, licensing, and commitment structures while informing the vendor renewal strategy.
β’ Maintain a working cadence with customer observability leadership.
β’ Generate assessment reports, target architecture, cost models, migration plans, executive summaries, and present findings to engineering leadership.
β’ Minimum of 8 years in SRE, DevOps, Observability, or Platform Engineering.
β’ At least 4 years of experience managing production observability platforms.
β’ Previous experience in a technical lead, staff, or principal-level role.
β’ Extensive experience in running metrics, logging, and tracing platforms at a large scale.
β’ Advanced production experience with Prometheus and Thanos, Cortex, or Mimir.
β’ Proficiency in PromQL, recording rules, remote write, and cardinality management.
β’ Practical experience with Datadog or a similar commercial platform.
β’ Production experience with CloudWatch, CloudWatch Logs, Container Insights, X-Ray or Application Signals, Amazon Managed Service for Prometheus, and Amazon Managed Grafana.
β’ Experience with OpenTelemetry Collector or ADOT, including pipeline design, processors, tail and head sampling, and multi-backend export.
β’ Background in designing and operating log pipelines using Vector, Fluent Bit, Logstash, or Firehose.
β’ Strong production experience with EKS or Kubernetes.
β’ High proficiency with Terraform.
β’ Experience in designing SLO-based alerting and integrating with incident management tools like PagerDuty.
β’ Ability to create defensible TCO models from usage data and pricing.
β’ Strong scripting skills in Python, Go, or Bash.
β’ Proven ability to assess unfamiliar environments and deliver recommendations rapidly.
β’ Capability to articulate technical, cost, and compliance tradeoffs to both technical and executive stakeholders.
β’ Experience with vendor migrations, dashboards and alerts as code, Grafana ecosystem tools, continuous profiling, data governance, analytics platforms, consumer-scale platforms, or consulting is a plus.
β’ Certifications such as Prometheus Certified Associate, OpenTelemetry Certified Associate, AWS Certified DevOps Engineer β Professional, AWS Certified Solutions Architect β Professional, or CKA are highly valued.
β’ EverOps offers remote hiring in the United States.
β’ Must possess legal authorization to work for any employer in the U.S.
β’ Fully remote workplace.
β’ Unlimited paid time off.
β’ Equity: Become a true owner of the company.
β’ 401K with company contribution.
β’ Sponsored healthcare.
β’ Access to training and certification programs.
Green Energy Venture AG
Abacus Group
EXP
Get handpicked remote jobs straight to your inbox weekly.