Site Reliability Engineer

Posted 2 days ago

This is a fully remote position, open to applicants in India.

📋 Description

• Deliver on-call production assistance, execute rapid triage, manage escalations, and ensure service restoration.

• Diagnose production and customer-related issues, spearhead incident troubleshooting, and facilitate root-cause analysis.

• Utilize AIOps-supported RCA, log clustering, timeline reconstruction, and incident summarizations.

• Take ownership and enhance the ELK/OpenSearch observability framework.

• Design and uphold observability across logs, metrics, and traces.

• Develop workflows for alert notifications, anomaly detection, and incident enrichment.

• Implement dynamic baselining, alert correlation, event suppression, impact forecasting, and automated context enrichment.

• Create automation, runbooks, and controlled self-healing processes with safety measures, rollback strategies, and audit trails.

• Link observability signals to runbooks, tickets, ChatOps actions, and human-approved recovery procedures.

• Enhance system reliability, performance, scalability, and resilience.

• Collaborate with engineering teams to ensure deployment safety, operational readiness, and production stability.

• Maintain runbooks, documentation, and knowledge bases.

• Assist in capacity planning, performance optimization, and SLO-based reliability practices.

• Implement security measures, access controls, and operational guardrails.

• Oversee AIOps implementation from use-case definition to production rollout, validation, adoption tracking, and ongoing adjustments.


⛳️ Requirements

• Minimum of 4 years of experience in SRE, AIOps, or production engineering within large-scale cloud environments.

• Strong proficiency in Linux/Unix, networking, distributed systems, and AWS/Azure/GCP.

• In-depth knowledge of ELK/OpenSearch, including Logstash/Beats, Elasticsearch index design, scaling and tuning, Kibana querying and debugging, dashboards, and alerts.

• Hands-on experience with logs, metrics, and tracing methodologies.

• Capability to prepare telemetry for AIOps implementation, involving tagging, normalization, correlation keys, service mapping, and event metadata.

• Understanding of incident management, root-cause analysis, SLOs, and operational best practices.

• Practical experience with AIOps implementation, including anomaly detection, event correlation, signal enrichment, noise suppression, incident summaries, and governed remediation.

• Ability to implement integrations across observability tools, ITSM/ticketing systems, ChatOps, CMDB/runbook repositories, and automation platforms.

• Experience in assessing AIOps effectiveness using operational KPIs.

• Familiarity with Infrastructure as Code tools such as Terraform, Ansible, or similar.

• Preferred certifications include AWS DevOps Engineer, AWS ML Specialty, Google Cloud DevOps Engineer, Azure DevOps Engineer, Kubernetes/CKA, or relevant certifications in AI/ML, AIOps, observability, or cloud automation.

• Willingness to work shifts – EMEA shifts.


🏝️ Benefits

• Remote-first organization.

• Flexible remote work arrangement.

• Access to Employee Resource Groups.

• Participate in "Coffee with Mark" sessions.

• Engage in Microsoft Teams communities centered on wellness, art, pets, family, and parenting.

• Attend special guest discussions on topics affecting employees.

People also viewed

FourEnergy GmbH13 hours ago

Senior DevOps Engineer – Operations

DE flagGermany OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
ICF15 hours ago

Lead DevOps Engineer

US flagVirginia OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$131.3k – $223.1k/year
ApplyView job
Mastercam19 hours ago

DevSecOps Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
C&S Informática1 day ago

DevOps Engineer – Freelance/Contract, Mid-Level/Senior

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Convene1 day ago

Support and Deployment Engineer

SA flagSaudi Arabia OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Verity Group1 day ago

SRE Engineer

BR flagBrazil OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers