
Senior Observability Analyst
Posted 2 days ago

Posted 2 days ago
This is a fully remote position, open to applicants in California.
• Spearhead the advancement of the organization's observability strategy, utilizing Datadog as the main enterprise platform.
• Act as the technical and functional subject-matter authority for the Datadog platform.
• Design, develop, and improve solutions utilizing APM, Infrastructure Monitoring, Logs Management, Dashboards, RUM, Synthetic Monitoring, and Continuous Testing.
• Establish enterprise observability standards for applications, APIs, microservices, and cloud workloads.
• Create and maintain executive, operational, and analytical dashboards.
• Formulate proactive monitoring strategies based on SLIs, SLOs, SLAs, and business indicators.
• Develop, review, and optimize intelligent monitors and alerts to minimize operational noise and reduce false positives.
• Assist with application instrumentation using OpenTelemetry and native Datadog integrations.
• Lead root cause investigations of critical incidents and propose preventive measures.
• Identify automation opportunities, failure prediction, and self-healing solutions.
• Create training materials, playbooks, standards, and documentation pertaining to Datadog.
• Execute capacity, performance, availability, and end-user experience analyses.
• Foster a culture of observability and operational excellence.
• Collaborate with Engineering, Architecture, Development, SRE, and Operations teams.
• Bachelor’s degree in Computer Science, Engineering, Information Systems, or a related discipline.
• Extensive experience with Datadog, including platform implementation, administration, and enhancement.
• Practical experience with Datadog APM, Infrastructure Monitoring, Logs Management, Dashboards, Monitors, and Monitoring Service Catalog.
• Advanced understanding of logs, metrics, and distributed tracing.
• Experience with observability in APIs and microservices architectures.
• Experience working in AWS environments.
• Hands-on experience in troubleshooting and investigating complex incidents.
• Knowledge of OpenTelemetry and application instrumentation techniques.
• Familiarity with automation using Python, Shell scripting, or PowerShell.
• Experience with Kubernetes, Docker, and cloud-native ecosystems.
• Strong understanding of system availability, performance, scalability, and reliability.
• Ability to interpret technical indicators in terms of business impact.
• Exceptional communication skills and the capacity to collaborate with multiple stakeholders.
• Datadog Certified Associate certification or higher is preferred.
• Experience with Site Reliability Engineering (SRE) practices is preferred.
• Experience in implementing observability strategies for large-scale distributed environments is preferred.
• Knowledge of CI/CD and DevSecOps practices is preferred.
• Experience with automated incident response and self-healing processes is preferred.
• Familiarity with operational governance, incident management, Problem Management, and ITIL processes is preferred.
• Experience with Dynatrace, Grafana, Prometheus, Elastic Stack, or Zabbix is preferred.
• Remote work opportunities.
• An inclusive, purpose-driven work environment.
• Flexibility to balance your career with personal commitments and interests.
• Well-being initiatives to support employees.
• Career development opportunities and professional experiences.
• Recognized as a Great Place To Work™ in 24 countries.
• Certified as an International Top Employer.
Tenet Healthcare
Early Childhood Educators
Coinbase
Get handpicked remote jobs straight to your inbox weekly.