
Senior Engineer β Observability Platform
Posted Aug 18

Posted Aug 18
This is a fully remote position, open to applicants in Louisiana, +4 more states.
β’ Create tailored software solutions for the observability platform utilizing Java Spring Boot, Python, and Go.
β’ Design, construct, and manage essential services for the observability platform.
β’ Drive organization-wide implementation of OpenTelemetry, encompassing client libraries, semantic conventions, instrumentation patterns, and Collector/agent strategies.
β’ Design, architect, and scale high-throughput, fault-tolerant telemetry pipelines for logs, metrics, and traces.
β’ Develop self-service observability features to facilitate onboarding, troubleshooting, and adoption.
β’ Implement comprehensive monitoring, SLOs, health checks, and alerting mechanisms for the observability platform.
β’ Set reliability benchmarks, error budgets, and incident response protocols in collaboration with SRE, Platform, and Cloud teams.
β’ Engage in on-call rotations, leading incident mitigation, root-cause analysis, and post-incident evaluations.
β’ Streamline operational workflows through tooling, CI/CD enhancements, and platform automation.
β’ Ensure the security of telemetry pipelines utilizing mTLS, secrets management, and zero-trust design methodologies.
β’ Create and maintain technical documentation, standards, and best practices.
β’ Collect requirements, influence roadmap prioritization, and implement platform enhancements.
β’ Provide technical leadership by mentoring, conducting design reviews, offering architectural guidance, and fostering cross-team collaboration.
β’ Minimum of 5 years of experience in Software Engineering.
β’ At least 3 years of experience with observability practices, including SLIs/SLOs/SLAs, alerting, and incident management.
β’ A minimum of 3 years developing production-grade backend services in Go and/or Java.
β’ At least 3 years of experience implementing and managing OpenTelemetry, including OTLP, semantic conventions, and instrumentation patterns.
β’ A minimum of 3 years working with cloud-native and containerized platforms, including Docker, Kubernetes, and Argo CD.
β’ At least 3 years of experience with public cloud platforms such as AWS, GCP, or Azure.
β’ Minimum of 2 years designing and scaling distributed, high-volume data pipelines.
β’ At least 2 years of experience with Infrastructure as Code tools like Terraform or CloudFormation.
β’ A minimum of 2 years with Helmcharts, Kustomize, or similar tools.
β’ At least 2 years working with Grafana OSS or comparable observability backends such as Grafana, Loki, Tempo, or Mimir.
β’ Minimum of 2 years with relational databases like PostgreSQL or MySQL.
β’ Bachelorβs degree from an accredited institution or equivalent work experience.
β’ Alternatively, a high school diploma plus 4 years of relevant experience.
β’ Preferred experience with service meshes, Envoy, Istio, commercial observability platforms, on-call scheduling tools, streaming/data platforms, time-series/NoSQL/analytical databases, cost optimization, chaos engineering, security-aware platform design, mentoring, 24x7 production systems, regulated environments, and cross-team technical communication.
β’ CVS Health bonus, commission or short-term incentive program in addition to base pay.
β’ Medical coverage.
β’ Dental coverage.
β’ Vision coverage.
β’ Paid time off.
β’ Retirement savings options.
β’ Wellness programs.
β’ Additional resources supporting physical, emotional, and financial well-being.
β’ Comprehensive benefits package for colleagues and their families, based on eligibility.
NVIDIA
FXC Intelligence
Miratech
FXC Intelligence
Get handpicked remote jobs straight to your inbox weekly.