
Lead Engineer – Observability Platform
Posted Aug 18

Posted Aug 18
This is a fully remote position, open to applicants in Louisiana, +4 more states.
• Design, develop, and manage core observability platform services utilizing Go, Python, and Java (Spring Boot).
• Spearhead the enterprise-wide implementation of OpenTelemetry, encompassing client libraries, semantic conventions, instrumentation patterns, and Collector/agent strategy.
• Architect and enhance high-throughput, fault-tolerant telemetry pipelines for logs, metrics, and traces.
• Create self-service observability capabilities to facilitate onboarding, troubleshooting, and overall adoption.
• Execute comprehensive monitoring, SLOs, health checks, and alerting for the observability platform.
• Collaborate with SRE, Platform, and Cloud teams to set reliability standards, error budgets, and incident response methodologies.
• Engage in on-call rotations and lead incident response, root-cause analysis, and post-incident evaluations.
• Streamline operational workflows through tooling, CI/CD enhancements, and platform automation.
• Safeguard telemetry pipelines using mTLS, secrets management, and zero-trust design principles.
• Generate and maintain technical documentation, standards, and best practices.
• Collect requirements from internal engineering teams, shape roadmap prioritization, and implement platform enhancements.
• Provide technical leadership through mentorship, design reviews, architectural guidance, and cross-team collaboration.
• Over 7 years of experience in Software Engineering.
• More than 5 years of experience with observability practices, including SLIs/SLOs/SLAs, alerting, and incident management.
• At least 5 years of experience building production-grade backend services in Go and/or Java.
• More than 5 years implementing and managing OpenTelemetry, including OTLP, semantic conventions, and instrumentation patterns.
• Over 5 years of experience with cloud-native and containerized environments, including Docker, Kubernetes, and Argo CD.
• More than 5 years working with public cloud services such as AWS, GCP, or Azure.
• At least 3 years of experience designing and scaling distributed, high-volume data pipelines.
• More than 3 years of experience with Helm charts and Kustomize.
• At least 3 years working with Grafana OSS or similar observability backends, including Grafana, Loki, Tempo, or Mimir.
• Over 3 years of experience with Infrastructure as Code tools such as Terraform or CloudFormation.
• More than 3 years of experience with relational databases like PostgreSQL or MySQL.
• Bachelor's degree from an accredited institution or equivalent work experience; alternatively, a high school diploma plus 4 years of relevant experience.
• Familiarity with service meshes and networking technologies such as Envoy and Istio.
• Experience with commercial observability platforms like Datadog, New Relic, or AppDynamics.
• Experience using on-call scheduling tools such as OpsGenie, PagerDuty, or GoAlert.
• Knowledge of streaming and data platforms such as Kafka or Pulsar.
• Familiarity with time-series, NoSQL, or analytical databases like ClickHouse, Bigtable, or Cassandra.
• Experience with cost optimization and capacity planning for large-scale telemetry systems.
• Background in chaos engineering, resiliency testing, or fault injection.
• Experience in security-focused platform design, including secure service-to-service communication.
• Proven track record of mentoring senior engineers and influencing platform standards across organizations.
• Strong operational experience in supporting 24x7 production systems, including on-call duties.
• Excellent technical communication and collaboration skills across teams.
• Experience operating within regulated or compliance-intensive sectors such as healthcare or finance.
• CVS Health bonus, commission, or short-term incentive program in addition to the base salary range.
• Comprehensive medical coverage.
• Dental coverage.
• Vision coverage.
• Paid time off.
• Retirement savings options.
• Wellness programs.
• Additional resources promoting physical, emotional, and financial well-being, subject to eligibility.
• Comprehensive benefits package available for full-time employees and their families.
Bixal
Ibrowse Consultoria e Informática
Cisco
PowerSchool
Get handpicked remote jobs straight to your inbox weekly.