
Senior Monitoring and Observability Engineer
Posted Sep 1

Posted Sep 1
This is a fully remote position, open to applicants in United States.
β’ Develop and oversee telemetry pipelines from start to finish utilizing Dynatrace, Datadog, and Splunk.
β’ Implement instrumentation for applications and infrastructure to capture metrics, logs, and traces based on OpenTelemetry standards.
β’ Create observability solutions for microservices, containers, and cloud-native workloads within AWS/GovCloud environments.
β’ Construct dashboards and enable golden-signal monitoring with correlation among metrics, events, logs, and traces.
β’ Supervise observability platforms and address issues related to collectors, agents, and pipelines.
β’ Establish and refine production alerts with defined ownership, severity levels, and links to runbooks.
β’ Participate in on-call rotations and engage in incident response bridges, employing telemetry-driven triage and root-cause analysis.
β’ Integrate alert routing, deduplication, and maintenance-window management with ServiceNow and paging tools.
β’ Take part in post-incident reviews and report on trends related to MTTD/MTTA.
β’ Deploy and configure observability agents and collectors utilizing Ansible, Ansible Tower, and AAP.
β’ Develop GitLab CI/CD pipelines for observability-as-code and transition existing Jenkins jobs to GitLab.
β’ Collaborate with platform engineering to instrument infrastructure provisioned via Terraform.
β’ Enhance versioned, peer-reviewed configuration-as-code for pipelines, dashboards, and alerts.
β’ Assist in managing tagging, lifecycle governance, retention policies, telemetry costs, ingest volumes, cardinality, data masking, auditing, and continuous monitoring.
β’ Maintain architecture documentation and runbooks while working within a SAFe framework using ServiceNow, Jira, and Confluence.
β’ U.S. Citizenship is required.
β’ Must be able to obtain and maintain the necessary Public Trust level clearance.
β’ Bachelor's Degree with 8 years of relevant experience, a Master's Degree with 6 years, or a High School diploma (or equivalent) with 12 years of experience.
β’ At least 4 years of hands-on experience with enterprise observability, monitoring, and alerting platforms.
β’ Practical experience with Dynatrace, Datadog, and/or Splunk, including tasks related to instrumentation, dashboard creation, and alert design.
β’ Proficiency in metrics, logs, and traces, along with experience in OpenTelemetry-based instrumentation.
β’ Experience working within AWS and/or AWS GovCloud.
β’ Familiarity with containerized and cloud-native workloads.
β’ Experience with Ansible / Ansible Tower / AAP for automated deployment of agents and configurations.
β’ Knowledge of GitLab CI/CD and Jenkins for developing observability-as-code pipelines.
β’ Strong scripting and automation skills in Python, Bash, or Go.
β’ Possible eligibility for overtime compensation.
β’ Possible eligibility for shift differentials.
β’ Possible eligibility for a discretionary bonus.
Northrop Grumman
Thermo Systems
Thermo Systems
Thermo Systems
Get handpicked remote jobs straight to your inbox weekly.