Remotery

Reliability Operations Engineer

Posted Jun 24

This is a fully remote position, open to applicants in Malaysia.

📋 Description

• Oversee incident investigations during the daytime hours of your region, offering timely updates, escalating as necessary, and assisting senior engineers in the response efforts.

• Address escalations from Tier 1 support by utilizing established runbooks, metrics, logs, and diagnostics to resolve issues or escalate to Tier 3 when appropriate.

• Revise runbooks and operational documentation based on new incidents, findings, and feedback, ensuring clarity and consistency throughout all procedures.

• Execute existing automations and collaborate with senior team members to improve tools and scripts that optimize troubleshooting and remediation activities.

• Utilize observability tools such as Grafana/Prometheus, GCP Monitoring, and OpenTelemetry to analyze metrics, logs, and traces, aiding in the identification of anomalies and the validation of system performance.

• Deliver concise and accurate updates during incidents, ensuring that information is communicated to the appropriate engineering and SRE contacts and facilitating structured incident coordination.

• Engage in discussions regarding root causes, share operational insights, and contribute to process improvements that enhance system stability and maintainability.

• Participate in a shared weekend on-call rotation to ensure operational coverage for production systems, responding to incidents and escalations as necessary and coordinating with engineering teams when issues arise.

• Actively enhance workflows, embrace best practices, and establish the foundation of the Reliability Operations function as it develops.


⛳️ Requirements

• Bachelor’s degree in Computer Science, Information Technology, Engineering, or equivalent hands-on experience.

• 2–4 years of experience in Reliability Operations, Site Reliability Engineering, DevOps, IT Operations, or a related technical support role.

• Experience in participating in Tier 1 or Tier 2 investigations, including log analysis, basic triage, and structured escalation processes.

• Familiarity with operational environments that support distributed or cloud-based systems.

• Involvement in incident response workflows and/or on-call rotations.

• Proficient in Linux, including system navigation, log review, and basic diagnostics.

• Experience in using and contributing to runbooks and operational workflows.

• Capability to interpret metrics, logs, and traces using tools such as Grafana/Prometheus, Google Cloud Monitoring, and OpenTelemetry.

• Knowledge of cloud platforms, preferably Google Cloud Platform (GCP).

• Ability to adhere to documented remediation procedures, demonstrating sound judgment regarding escalation.

• Understanding of CI/CD pipelines and the impact of application deployments on runtime behavior.

• Familiarity with Jira or similar ticketing systems.

• Strong and effective communicator, particularly when providing updates during time-sensitive operational situations.

• Calm and organized approach to troubleshooting and prioritization.

• Collaborative mindset, effectively working with senior operations engineers, product teams, and SREs.

• Strong sense of ownership and accountability for operational responsibilities.


🏝️ Benefits

• Continuous operational coverage

• Weekend on-call rotation shared across the Reliability Operations team

People also viewed

Ontrac Solutions2 days ago

Site Reliability Engineer

PK flagPakistan OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
CyberSheath2 days ago

Cloud Operations Engineer

US flagVirginia OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$110k – $127k/year
ApplyView job
Ontrac Solutions2 days ago

Site Reliability Engineer

IN flagIndia OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
NVIDIA2 days ago

Service Reliability Engineer

US flagTexas OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$168k – $333.5k/year
ApplyView job
Nagarro2 days ago

Senior Site Reliability Engineer, AWS Cloud

RO flagRomania OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Capgemini2 days ago

Senior DevOps Engineer

UA flagUkraine OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers