
Senior Automation, Observability Engineer
Posted Sep 17

Posted Sep 17
This is a fully remote position, open to applicants in United States.
β’ Design, implement, and sustain enterprise monitoring and observability solutions.
β’ Develop and manage Grafana dashboards, alerts, and visualizations.
β’ Oversee monitoring of infrastructure, applications, middleware, IoT services, and enterprise telemetry using IBM Instana, Grafana, SolarWinds, and related tools.
β’ Configure and manage data collection with Telegraf, Prometheus, and monitoring agents.
β’ Analyze metrics, logs, traces, events, and telemetry to pinpoint bottlenecks and service degradation.
β’ Support monitoring initiatives regarding SLOs, SLAs, and operational health.
β’ Conduct root cause analysis and troubleshoot infrastructure and application challenges.
β’ Assist in onboarding, monitoring, and operational management of FOAK services and enterprise applications.
β’ Configure, validate, and troubleshoot ELT integrations and telemetry pipelines.
β’ Oversee log ingestion, event correlation, and telemetry data quality.
β’ Monitor Linux and Windows servers, VMware, Citrix VDI, DNS, proxy services, middleware, integration services, enterprise applications, and IoT platforms.
β’ Investigate performance issues, recurring alerts, infrastructure anomalies, capacity, availability, and service health.
β’ Support platform upgrades, maintenance, and operational readiness reviews.
β’ Configure and maintain InfluxDB, including retention policies, performance tuning, and capacity planning.
β’ Develop operational dashboards and reports to provide performance insights.
β’ Acknowledge, investigate, troubleshoot, and resolve incidents.
β’ Coordinate incident resolutions with Infrastructure, Network, Cloud, Security, Application, and Service Delivery teams.
β’ Participate in major incident bridges, disaster recovery exercises, and provide 24x7 operational support.
β’ Adhere to escalation procedures, SOPs, runbooks, and ITIL processes; support Problem Management and RCA documentation.
β’ Administer IBM Instana environments and assist in APM configuration, alerts, baselines, and thresholds.
β’ Develop automation scripts using Python, PowerShell, Bash, and VBScript for operational tasks, monitoring deployments, remediation, and event-driven workflows.
β’ Implement Ansible and Puppet for infrastructure automation, server provisioning, configuration management, agent deployment, playbooks, and pipelines.
β’ Maintain SOPs, runbooks, monitoring procedures, escalation matrices, and observability documentation.
β’ Engage in knowledge-transfer sessions, service onboarding, service transition, migration, and continuous improvement initiatives.
β’ Proficient in Grafana, IBM Instana, SolarWinds, Telegraf, Prometheus, InfluxDB, OpenTelemetry, Grafana Alloy, APM monitoring, event management, alert management, observability concepts, and SLO/SLA monitoring.
β’ Experience with FOAK support and Enterprise Logging & Telemetry (ELT).
β’ Knowledge of log aggregation and correlation, telemetry data analysis, event correlation, application onboarding, and monitoring standards/observability frameworks.
β’ Familiarity with VMware, Linux administration, Windows Server, Citrix VDI, DNS services, proxy services, middleware technologies, and infrastructure performance monitoring.
β’ Proficient in Python, PowerShell, Linux shell scripting, and VBScript.
β’ Experience with Ansible, Puppet, webhooks, and Infrastructure as Code (IaC).
β’ Knowledge of ServiceNow, incident management, problem management, change management, ITIL Framework, and major incident management.
β’ Preferred: 7 to 10+ years of experience in Monitoring, Observability, Infrastructure Operations, SRE, or Platform Engineering.
β’ Experience in supporting large-scale enterprise environments and 24x7 operations.
β’ Hands-on experience with Grafana, Instana, SolarWinds, Telegraf, Prometheus, and InfluxDB.
β’ Experience supporting VMware, Citrix, middleware, enterprise applications, and cloud monitoring platforms.
β’ Familiarity with FOAK applications and ELT platforms.
β’ Strong skills in troubleshooting, RCA, incident management, and operational support.
β’ Experience with automation frameworks and Infrastructure as Code (preferably Ansible).
β’ Experience integrating observability platforms with enterprise automation solutions.
β’ Nice-to-have: Docker, Kubernetes, AWS, Microsoft Azure, Google Cloud Platform, Jenkins, GitHub Actions, GitLab CI/CD, REST APIs, microservices monitoring, and DevOps/SRE practices.
β’ Unlimited Paid Days Off.
β’ Three health plan options.
β’ 401k with company match.
β’ Dental, vision, short-term disability, long-term disability, life, and AD&D coverage.
β’ Flexible spending accounts.
β’ Family Forming Benefit including fertility coverage and adoption/surrogacy reimbursement.
β’ Paid childbearing and paternal leave.
β’ Education Reimbursement.
β’ Student Loan Assistance or 529 College Funding.
β’ Sabbatical leave.
β’ Wellness program.
β’ Flexible work schedule.
β’ Annual bonus plan based on company and individual performance.
β’ Equity grant under the Associate Equity Appreciation Program.
β’ Option to work from home or in Ensono offices when not required on a client site.
Mercor
RTX
Expel
Qualus
Get handpicked remote jobs straight to your inbox weekly.