
Observability Technical Lead
Posted Sep 9

Posted Sep 9
This is a fully remote position, open to applicants in Florida.
• Act as the lead Subject Matter Expert (SME) for enterprise observability, establishing architectures, standards, integration patterns, and reusable solutions.
• Architect and enhance observability across cloud systems, on-premises infrastructure, networks, Kubernetes, containers, applications, APIs, databases, middleware, and enterprise platforms.
• Standardize metrics, logs, traces, events, topology, dashboards, alerting, instrumentation, and service health.
• Provide technical oversight for tools such as LogicMonitor, Splunk, OpenTelemetry, Grafana, Prometheus, Tempo, VictoriaMetrics, Loki, and other related technologies.
• Spearhead the adoption of OpenTelemetry, encompassing instrumentation, collectors, telemetry pipelines, distributed tracing, context propagation, and vendor-neutral standards.
• Design telemetry pipelines to route metrics, logs, and traces across various platforms.
• Create API-driven and Observability-as-Code functionalities for onboarding, configuration, validation, and lifecycle management.
• Establish governance surrounding RBAC, telemetry standards, alerting, retention, integrations, configuration management, and data lifecycle management.
• Enhance observability coverage, telemetry quality, scalability, reliability, performance, and cost-effectiveness.
• Propel automation and self-service initiatives through APIs, Infrastructure-as-Code, CI/CD, GitOps, and reusable observability patterns.
• Integrate Edwin AI and AIOps features for anomaly detection, event correlation, root-cause analysis, investigation, and operational intelligence.
• Connect telemetry with operational context to ensure comprehensive visibility.
• Set reliability benchmarks including SLIs, SLOs, error budgets, monitoring coverage, alert quality, MTTD, and MTTR.
• Mitigate alert fatigue using intelligent correlation, dynamic thresholds, suppression, automation, and event-management techniques.
• Assess eBPF, continuous profiling, Kubernetes observability, dependency mapping, and auto-instrumentation methods.
• Transition operations from reactive monitoring to proactive and predictive management.
• Collaborate with engineering and operations teams spanning North America, Europe, and Asia.
• Lead architecture reviews, platform assessments, workshops, and technical working sessions.
• Shape observability strategy and standards across teams without possessing direct authority.
• Mentor engineers and advocate for observability, Site Reliability Engineering (SRE), and reliability best practices.
• Evaluate emerging technologies and suggest adoption based on interoperability, scalability, business value, and cost considerations.
• Convert strategy into standards, reference architectures, and reusable implementation patterns.
• Bachelor's Degree.
• Over 7 years of experience in observability, monitoring, Application Performance Management (APM), SRE, DevOps, platform engineering, or a related field.
• More than 7 years of experience in designing, implementing, operating, and maintaining enterprise-scale observability platforms utilizing observability-as-code, monitoring-as-code, infrastructure-as-code, GitOps, and CI/CD practices.
• Proficient with AWS, Azure, GCP, and large-scale Kubernetes ecosystems.
• Experience in designing OpenTelemetry Collector architectures and telemetry pipelines.
• Expertise in Grafana, Tempo, Loki, and Prometheus-compatible platforms.
• Experience with VictoriaMetrics or comparable large-scale time-series databases.
• Familiarity with eBPF, continuous profiling, auto-instrumentation, and cloud-native telemetry.
• Experience integrating observability platforms with ServiceNow, ITSM, CMDB, incident management, and automation tools.
• Understanding of SRE practices including SLIs, SLOs, error budgets, and incident management.
• Experience in enabling developer self-service and internal developer platform integrations.
• Knowledge of RBAC, secrets management, governance, compliance, and telemetry data protection.
• Flexible scheduling to accommodate global collaboration.
• Short-term cash incentives, contingent upon plan requirements.
• Medical, Dental, and Vision coverage.
• Wellness incentives.
• Retirement benefits.
• Paid vacation days, up to 15 days.
• Paid sick days, up to 5 days.
• Paid personal leave, up to 5 days.
• Paid holidays, up to 13 days.
• Leave for birth and adoption.
• Parental leave.
• Family and medical leave.
• Bereavement leave.
• Jury duty leave.
• Military leave.
• Option to purchase additional vacation.
• Short-term and long-term disability coverage.
• Life Insurance and Accidental Death and Dismemberment coverage.
• Health Savings Account.
• Health Care Spending Account.
• Dependent Care Spending Account.
• Tuition Assistance.
LMI
Sourcegraph
Verra Mobility
Cint
Get handpicked remote jobs straight to your inbox weekly.