Observability Technical Lead

atCarrierRemoteUS flagFloridaFull-timeFull-stack EngineerSenior$96k – $192k/year

Posted Sep 9

This is a fully remote position, open to applicants in Florida.

📋 Description

• Act as the lead Subject Matter Expert (SME) for enterprise observability, establishing architectures, standards, integration patterns, and reusable solutions.

• Architect and enhance observability across cloud systems, on-premises infrastructure, networks, Kubernetes, containers, applications, APIs, databases, middleware, and enterprise platforms.

• Standardize metrics, logs, traces, events, topology, dashboards, alerting, instrumentation, and service health.

• Provide technical oversight for tools such as LogicMonitor, Splunk, OpenTelemetry, Grafana, Prometheus, Tempo, VictoriaMetrics, Loki, and other related technologies.

• Spearhead the adoption of OpenTelemetry, encompassing instrumentation, collectors, telemetry pipelines, distributed tracing, context propagation, and vendor-neutral standards.

• Design telemetry pipelines to route metrics, logs, and traces across various platforms.

• Create API-driven and Observability-as-Code functionalities for onboarding, configuration, validation, and lifecycle management.

• Establish governance surrounding RBAC, telemetry standards, alerting, retention, integrations, configuration management, and data lifecycle management.

• Enhance observability coverage, telemetry quality, scalability, reliability, performance, and cost-effectiveness.

• Propel automation and self-service initiatives through APIs, Infrastructure-as-Code, CI/CD, GitOps, and reusable observability patterns.

• Integrate Edwin AI and AIOps features for anomaly detection, event correlation, root-cause analysis, investigation, and operational intelligence.

• Connect telemetry with operational context to ensure comprehensive visibility.

• Set reliability benchmarks including SLIs, SLOs, error budgets, monitoring coverage, alert quality, MTTD, and MTTR.

• Mitigate alert fatigue using intelligent correlation, dynamic thresholds, suppression, automation, and event-management techniques.

• Assess eBPF, continuous profiling, Kubernetes observability, dependency mapping, and auto-instrumentation methods.

• Transition operations from reactive monitoring to proactive and predictive management.

• Collaborate with engineering and operations teams spanning North America, Europe, and Asia.

• Lead architecture reviews, platform assessments, workshops, and technical working sessions.

• Shape observability strategy and standards across teams without possessing direct authority.

• Mentor engineers and advocate for observability, Site Reliability Engineering (SRE), and reliability best practices.

• Evaluate emerging technologies and suggest adoption based on interoperability, scalability, business value, and cost considerations.

• Convert strategy into standards, reference architectures, and reusable implementation patterns.


⛳️ Requirements

• Bachelor's Degree.

• Over 7 years of experience in observability, monitoring, Application Performance Management (APM), SRE, DevOps, platform engineering, or a related field.

• More than 7 years of experience in designing, implementing, operating, and maintaining enterprise-scale observability platforms utilizing observability-as-code, monitoring-as-code, infrastructure-as-code, GitOps, and CI/CD practices.

• Proficient with AWS, Azure, GCP, and large-scale Kubernetes ecosystems.

• Experience in designing OpenTelemetry Collector architectures and telemetry pipelines.

• Expertise in Grafana, Tempo, Loki, and Prometheus-compatible platforms.

• Experience with VictoriaMetrics or comparable large-scale time-series databases.

• Familiarity with eBPF, continuous profiling, auto-instrumentation, and cloud-native telemetry.

• Experience integrating observability platforms with ServiceNow, ITSM, CMDB, incident management, and automation tools.

• Understanding of SRE practices including SLIs, SLOs, error budgets, and incident management.

• Experience in enabling developer self-service and internal developer platform integrations.

• Knowledge of RBAC, secrets management, governance, compliance, and telemetry data protection.

• Flexible scheduling to accommodate global collaboration.


🏝️ Benefits

• Short-term cash incentives, contingent upon plan requirements.

• Medical, Dental, and Vision coverage.

• Wellness incentives.

• Retirement benefits.

• Paid vacation days, up to 15 days.

• Paid sick days, up to 5 days.

• Paid personal leave, up to 5 days.

• Paid holidays, up to 13 days.

• Leave for birth and adoption.

• Parental leave.

• Family and medical leave.

• Bereavement leave.

• Jury duty leave.

• Military leave.

• Option to purchase additional vacation.

• Short-term and long-term disability coverage.

• Life Insurance and Accidental Death and Dismemberment coverage.

• Health Savings Account.

• Health Care Spending Account.

• Dependent Care Spending Account.

• Tuition Assistance.

People also viewed

LMI10 hours ago

Full Stack Developer

US flagUnited States OnlyFull-timeFull-stack Engineer$122.2k – $211.3k/year
ApplyView job
Sourcegraph10 hours ago

Tech Lead – Code Plane

EuropeFull-timeFull-stack Engineer$144k – $192k/year
ApplyView job
Verra Mobility10 hours ago

Vice President, Software Engineering

US flagTexas OnlyFull-timeFull-stack Engineer
ApplyView job
Cint10 hours ago

Staff Software Engineer – DSM Team

ES flagSpain OnlyFull-timeFull-stack Engineer
ApplyView job
CI&T10 hours ago

Mid Level FullStack Developer – .NET, React

BR flagBrazil OnlyFull-timeFull-stack EngineerR$1 – R$2/month
ApplyView job
Akamai Technologies10 hours ago

Senior Software Engineer

PL flagPoland OnlyFull-timeFull-stack Engineer
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers