
Application Observability Engineer
Posted Sep 16

Posted Sep 16
This is a fully remote position, open to applicants in Arizona, +17 more states.
• Oversee and enhance the enterprise observability platform, which includes Grafana, Grafana Tempo, Dynatrace, Splunk, and syslog.
• Track system health and performance metrics to proactively detect and address potential challenges.
• Diagnose and resolve intricate issues concerning application performance, reliability, and functionality.
• Act as a subject matter expert and primary escalation point for observability and performance-related matters.
• Develop and sustain dashboards, alerts, and reports to facilitate operational visibility.
• Collaborate with network engineering and automation teams to ensure the smooth integration and functioning of monitoring systems.
• Lead root cause analysis for critical incidents and take part in post-mortem evaluations.
• Guide junior engineers on observability tools and troubleshooting techniques.
• Document processes, configurations, and troubleshooting procedures to promote knowledge sharing and operational efficiency.
• Engage in on-call rotations and incident response as necessary.
• Carry out other job-related tasks as assigned.
• Bachelor's degree in Computer Science, Information Technology, or a related field (or equivalent experience).
• A minimum of 5 years of experience in observability, operations, or related infrastructure roles.
• Practical experience with Grafana, Grafana Tempo, Dynatrace, Splunk, and syslog.
• Strong expertise with Linux systems and command-line tools.
• Comprehensive understanding of networking fundamentals (TCP/IP, DNS, routing, etc.).
• Experience with Kubernetes and containerized application environments.
• Proficiency in scripting languages (Python, Bash, or similar).
• Exceptional problem-solving and communication abilities.
• Capability to work independently in a remote environment.
• Willingness to collaborate as part of a team and also work autonomously.
• Must possess legal authorization to work in the U.S.
• Experience in telecommunications or large-scale distributed systems environments.
• Familiarity with distributed tracing and APM concepts.
• Understanding of CI/CD pipelines and DevOps methodologies.
• Experience with infrastructure-as-code tools (Terraform, Ansible).
• 80% of medical premiums covered by the company.
• Company-sponsored disability and life insurance.
• 401(k) with company matching.
• Discounted services within coverage areas.
• Competitive total compensation package.
• Locally owned, welcoming, and enjoyable work environment.
• Equal opportunity employment.
Mercor
RTX
Expel
Qualus
Get handpicked remote jobs straight to your inbox weekly.