
Incident Operation Engineer
Posted Aug 31

Posted Aug 31
This is a fully remote position, open to applicants in Romania.
• Oversee and react to real-time alerts, prioritize incidents, and assist in incident response efforts.
• Collaborate with Engineering, Incident Command, Customer Support, and Operations teams to enhance incident resolution efficiency.
• Evaluate incident impact, severity, and customer exposure using monitoring tools and system insights.
• Handle customer-facing communications, including incident notifications, status updates, and resolution reports.
• Administer and update public status pages, ensuring information is timely and accurate.
• Participate in post-incident reviews, root cause analysis processes, service level agreement reporting, and operational enhancements.
• Lead initiatives for automation and process optimization using technologies like Python or Kotlin.
• Assist in improving monitoring, observability, escalation processes, and operational tools.
• 7+ years of experience in Incident Management, Site Reliability Engineering (SRE), Technical Operations, Production Operations, or a related field.
• Experience in on-call and service level agreement (SLA) focused environments.
• In-depth knowledge of distributed systems, production environments, and service reliability.
• Practical experience with monitoring tools such as Datadog, Grafana, Prometheus, or comparable platforms.
• Familiarity with incident management tools like PagerDuty, Opsgenie, ServiceNow, or Rootly.
• Proficiency in programming with Python or Kotlin.
• Excellent communication skills with the ability to handle high-pressure situations and manage multiple priorities effectively.
• Comprehensive health and wellness benefits.
• Opportunities for professional development and growth.
• Flexible work arrangements and a supportive team environment.
Mercor
RTX
Expel
Qualus
Get handpicked remote jobs straight to your inbox weekly.