Remotery

Site Reliability Engineer – Level 3

Posted Jul 29

This is a fully remote position, open to applicants in India.

📋 Description

• Deliver on-call production support, ensuring swift triage, escalation management, and service recovery.

• Analyze production and customer-related issues, lead incident investigations, and facilitate prompt root cause analysis (RCA) with clear follow-up actions.

• Leverage AIOps-assisted RCA, log clustering, timeline reconstruction, and incident summarization to expedite diagnosis while verifying AI recommendations prior to execution.

• Take ownership of and enhance the observability stack, possessing thorough knowledge of ELK/OpenSearch (Elasticsearch, Logstash, Kibana) for log ingestion, indexing, querying, visualization, and alerting.

• Design and uphold observability across logs, metrics, and traces, guaranteeing actionable monitoring and optimal signal-to-noise ratios in alerts.

• Develop and refine workflows for alerting, anomaly detection, and incident enrichment to minimize noise and elevate accuracy.

• Execute AIOps use cases such as dynamic baselining, alert correlation, event suppression, impact prediction, and automated context enrichment.

• Create automation, runbooks, and controlled self-healing mechanisms with appropriate safeguards, rollback strategies, and traceability.

• Implement AIOps remediation patterns that link observability signals to runbooks, tickets, ChatOps actions, and human-approved recovery procedures.

• Propel enhancements in system reliability, performance, scalability, and resilience through engineering-led initiatives.

• Collaborate with engineering teams to enhance deployment safety, operational readiness, and production stability.

• Maintain comprehensive runbooks, documentation, and knowledge bases to boost on-call efficiency and knowledge sharing.

• Assist in capacity planning, performance optimization, and SLO-based reliability practices.

• Enforce security, access controls, and operational guardrails across systems and automation.

• Oversee AIOps implementation from use-case definition to production deployment, encompassing telemetry readiness, model/rule configuration, integration testing, operational validation, adoption tracking, and continuous tuning.


⛳️ Requirements

• 6+ years of experience in SRE, AIOps, or production engineering within large-scale cloud environments.

• Strong proficiency in Linux/Unix, networking, distributed systems, and cloud platforms (AWS/Azure/GCP).

• Expertise in ELK/OpenSearch, including: log ingestion (Logstash / Beats), Elasticsearch index design, scaling, and tuning; advanced Kibana querying and debugging; dashboards, alerts, and observability patterns for production systems.

• Practical experience with logs, metrics, and tracing.

• Capability to prepare telemetry for AIOps implementation, including tagging, normalization, correlation keys, service mapping, and high-quality event metadata.

• Robust understanding of incident management, RCA, SLOs, and operational best practices.

• Good grasp of AIOps concepts: anomaly detection, alert correlation, and intelligent alerting.

• Hands-on experience in implementing AIOps workflows, including event correlation, signal enrichment, noise suppression, automated incident summaries, and governed remediation.

• Ability to integrate AIOps across observability tools, ITSM/ticketing systems, ChatOps, CMDB/runbook repositories, and automation platforms.

• Experience in measuring AIOps effectiveness through operational KPIs such as alert-noise reduction, quicker MTTD/MTTR, RCA quality, automation adoption, repeat usage, and linkage to business impact.

• Familiarity with Infrastructure as Code tools such as Terraform, Ansible, or similar.

• Preferred certifications: AWS DevOps Engineer, AWS ML Specialty, Google Cloud DevOps Engineer, Azure DevOps Engineer, Kubernetes/CKA, or relevant AI/ML, AIOps, observability, or cloud automation certifications.


🏝️ Benefits

• Employee Resource Groups to promote diverse voices.

• Coffee with Mark sessions – Our employees have the opportunity to engage with our CEO on significant and sometimes challenging topics ranging from mental health to work-life balance and current events.

• Microsoft Teams communities centered on wellness, art, pets, family, parenting, and more.

• Occasional special guests to discuss issues that affect our employee community.

People also viewed

DATAGROUP1 day ago

DevOps Engineer

DE flagGermany OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Ambush1 day ago

DevOps Engineer

BR flagBrazil OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
DuoKey1 day ago

DevOps Engineer

MU flagMauritius OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
TEKsystems2 days ago

SRE – CloudOps, Practice Architect II

US flagIllinois OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$148.2k – $222.4k/year
ApplyView job
TEKsystems2 days ago

SRE CloudOps Practice Architect II

US flagTexas OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$148.2k – $222.4k/year
ApplyView job
Level Data2 days ago

Senior DevOps Engineer

US flagMassachusetts OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$120k – $135k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers