
Operations Intelligence Engineer
Posted Jul 30

Posted Jul 30
This is a fully remote position, open to applicants in Portugal.
• Coordinate and respond to service-impacting events effectively.
• Serve as the primary incident coordinator for incidents of low to medium complexity.
• Investigate incidents using various monitoring, logging, and observability tools.
• Examine alert patterns to find ways to enhance signal quality.
• Collaborate with service owners to increase monitoring efficiency and detection coverage.
• Assist in the administration and support of BigPanda and related observability tools.
• Engage in SIE retrospectives to pinpoint areas for operational enhancement.
• Create and update runbooks, knowledge articles, and operational procedures.
• Contribute to KPI reporting and efforts in operational analytics.
• Offer guidance and mentorship to Operations Intelligence Associates.
• Participate in after-hours incident response and on-call duties as necessary.
• Work alongside Engineering teams to resolve recurring operational challenges.
• Bachelor's degree or a comparable combination of technical education and relevant experience is preferred.
• 3–5 years of experience supporting enterprise production environments.
• Strong understanding of IP, TCP, UDP, SSL/TLS protocols, and other networking concepts is essential.
• Experience with Telephony, PBX phone systems, SIP/RTP, and WebRTC is preferred.
• Familiarity with common tools for validating network connectivity is required.
• Capability to document operational procedures, troubleshooting efforts, and incident response workflows.
• Self-motivated with an eagerness to learn new technologies and methodologies.
• Experience in supporting distributed cloud-based applications and services.
• Proficient in using monitoring, logging, or observability platforms.
• Knowledge of incident management processes is necessary.
• Familiarity with cloud environments such as AWS or Azure is required.
• Excellent troubleshooting and documentation skills are a must.
• Experience with operational platforms like ServiceNow, Jira, or similar tools is preferred.
• **Bonus Skills:**
• Exposure to cloud platforms including AWS, Azure, or GCP.
• Experience with observability tools such as Grafana, Prometheus, Kibana, Splunk, Datadog, Dynatrace, New Relic, LogicMonitor, or similar.
• Familiarity with event intelligence tools like BigPanda, PagerDuty, ServiceNow Event Management, Moogsoft, or similar.
• Experience in supporting Linux and Windows server environments.
• Knowledge of virtualization technologies such as VMware or Hyper-V.
• Familiarity with operational analytics, KPI reporting, or dashboard development.
• Experience in incident management or major incident response activities.
• We hire, promote, and reward employees based on their job performance, irrespective of race, color, creed, religion, sex, gender, marital status, national origin, ancestry, age, citizenship, physical or mental disability, sexual orientation, or any other protected status as outlined in our Code of Conduct (collectively referred to as “Protected Classes”). We maintain a strict no-discrimination policy in our workplace and are dedicated to providing reasonable accommodations for recognized disabilities or limitations as required by applicable laws. We are an equal opportunity employer and prioritize diversity within our organization. Discrimination based on race, religion, color, national origin, gender, sexual orientation, age, marital status, veteran status, or disability status is not tolerated.
Frontier Airlines
24-MAG
Varsity Brands
RTX
Get handpicked remote jobs straight to your inbox weekly.