
Expert Automation, Observability Engineer
Posted Sep 18

Posted Sep 18
This is a fully remote position, open to applicants in United States.
β’ Act as the strategic technical lead and architect for enterprise Observability, APM, and Telemetry ecosystems.
β’ Transition decentralized, reactive monitoring into a consolidated, automated, proactive observability framework.
β’ Design and oversee observability across metrics, logs, traces, and events.
β’ Direct FOAK implementations and transform new observability technologies into secure, repeatable, production-ready patterns.
β’ Establish standards for telemetry pipelines, data retention, high-cardinality controls, and observability cost management.
β’ Oversee SLIs, SLOs, and error budgets.
β’ Facilitate P1/P2 incident war rooms and perform evidence-based Root Cause Analysis.
β’ Minimize MTTD, MTTR, and alert noise through event correlation, dynamic thresholds, and dependency mapping.
β’ Promote Observability-as-Code and infrastructure automation.
β’ Automate deployment, configuration, and self-healing workflows for monitoring agents and telemetry collectors.
β’ Integrate observability platforms with ServiceNow, Netcool, and CI/CD pipelines.
β’ Design observability for Docker, Kubernetes, microservices, and multi-cloud environments.
β’ Correlate application APM telemetry with Kubernetes and infrastructure dependencies.
β’ Ensure telemetry pipelines are secure by design.
β’ Lead knowledge transfer programs, vendor transitions, and operational readiness handovers for global 24x7 teams.
β’ Guide cross-functional engineering teams and impact enterprise technology roadmaps.
β’ Over 12 years of comprehensive IT experience.
β’ At least 5 to 7 years serving as a Lead Architect, SRE, or Principal Observability Engineer in a large enterprise setting.
β’ Demonstrated success in migrating organizations from legacy monitoring systems to proactive, automated observability platforms.
β’ Practical expertise in developing scalable, secure telemetry pipelines and time-series databases.
β’ Significant experience leading FOAK rollouts and complex vendor/operations transition (KT) initiatives.
β’ Proficient with IBM Instana, Grafana, Prometheus, OpenTelemetry, Telegraf, and InfluxDB.
β’ Familiar with SolarWinds, Netcool, and Elastic/Splunk.
β’ Experience with Kubernetes, Docker, OpenShift, and cloud platforms such as AWS/Azure/GCP.
β’ Knowledgeable in Linux (RHEL), Windows Server, VMware, Citrix VDI, load balancers, and edge proxies.
β’ Skilled in Ansible, Terraform, Python, Bash, Webhooks, and CI/CD using GitHub Actions, GitLab, or Jenkins.
β’ Experience with ServiceNow, ITIL 4, and advanced Major Incident Management.
β’ Preferred certifications include CKA, Cloud Architect certification, or specific APM/Observability vendor certifications.
β’ Unlimited Paid Days Off.
β’ Three health plan options.
β’ 401k with company match.
β’ Access to dental, vision, short and long-term disability, life and AD&D coverage, and flexible spending accounts.
β’ Family Forming Benefit encompassing fertility coverage and adoption/surrogacy reimbursement.
β’ Paid parental leave for childbearing and paternal responsibilities.
β’ Education Reimbursement, Student Loan Assistance, or 529 College Funding.
β’ Sabbatical leave.
β’ Wellness program.
β’ Flexible work schedule.
β’ Annual bonus plan based on both company and individual performance.
β’ Equity grant under the Associate Equity Appreciation Program.
β’ Option to work remotely or in Ensono offices when not required at a client site.
Mercor
RTX
Expel
Qualus
Get handpicked remote jobs straight to your inbox weekly.