
Reliability Operations Engineer
Posted Jun 24

Posted Jun 24
This is a fully remote position, open to applicants in Malaysia.
• Oversee incident investigations during the daytime hours of your region, offering timely updates, escalating as necessary, and assisting senior engineers in the response efforts.
• Address escalations from Tier 1 support by utilizing established runbooks, metrics, logs, and diagnostics to resolve issues or escalate to Tier 3 when appropriate.
• Revise runbooks and operational documentation based on new incidents, findings, and feedback, ensuring clarity and consistency throughout all procedures.
• Execute existing automations and collaborate with senior team members to improve tools and scripts that optimize troubleshooting and remediation activities.
• Utilize observability tools such as Grafana/Prometheus, GCP Monitoring, and OpenTelemetry to analyze metrics, logs, and traces, aiding in the identification of anomalies and the validation of system performance.
• Deliver concise and accurate updates during incidents, ensuring that information is communicated to the appropriate engineering and SRE contacts and facilitating structured incident coordination.
• Engage in discussions regarding root causes, share operational insights, and contribute to process improvements that enhance system stability and maintainability.
• Participate in a shared weekend on-call rotation to ensure operational coverage for production systems, responding to incidents and escalations as necessary and coordinating with engineering teams when issues arise.
• Actively enhance workflows, embrace best practices, and establish the foundation of the Reliability Operations function as it develops.
• Bachelor’s degree in Computer Science, Information Technology, Engineering, or equivalent hands-on experience.
• 2–4 years of experience in Reliability Operations, Site Reliability Engineering, DevOps, IT Operations, or a related technical support role.
• Experience in participating in Tier 1 or Tier 2 investigations, including log analysis, basic triage, and structured escalation processes.
• Familiarity with operational environments that support distributed or cloud-based systems.
• Involvement in incident response workflows and/or on-call rotations.
• Proficient in Linux, including system navigation, log review, and basic diagnostics.
• Experience in using and contributing to runbooks and operational workflows.
• Capability to interpret metrics, logs, and traces using tools such as Grafana/Prometheus, Google Cloud Monitoring, and OpenTelemetry.
• Knowledge of cloud platforms, preferably Google Cloud Platform (GCP).
• Ability to adhere to documented remediation procedures, demonstrating sound judgment regarding escalation.
• Understanding of CI/CD pipelines and the impact of application deployments on runtime behavior.
• Familiarity with Jira or similar ticketing systems.
• Strong and effective communicator, particularly when providing updates during time-sensitive operational situations.
• Calm and organized approach to troubleshooting and prioritization.
• Collaborative mindset, effectively working with senior operations engineers, product teams, and SREs.
• Strong sense of ownership and accountability for operational responsibilities.
• Continuous operational coverage
• Weekend on-call rotation shared across the Reliability Operations team
Ontrac Solutions
CyberSheath
Ontrac Solutions
NVIDIA
Get handpicked remote jobs straight to your inbox weekly.