
Site Reliability Engineer – Level 3
Posted Jul 29

Posted Jul 29
This is a fully remote position, open to applicants in India.
• Deliver on-call production support, ensuring swift triage, escalation management, and service recovery.
• Analyze production and customer-related issues, lead incident investigations, and facilitate prompt root cause analysis (RCA) with clear follow-up actions.
• Leverage AIOps-assisted RCA, log clustering, timeline reconstruction, and incident summarization to expedite diagnosis while verifying AI recommendations prior to execution.
• Take ownership of and enhance the observability stack, possessing thorough knowledge of ELK/OpenSearch (Elasticsearch, Logstash, Kibana) for log ingestion, indexing, querying, visualization, and alerting.
• Design and uphold observability across logs, metrics, and traces, guaranteeing actionable monitoring and optimal signal-to-noise ratios in alerts.
• Develop and refine workflows for alerting, anomaly detection, and incident enrichment to minimize noise and elevate accuracy.
• Execute AIOps use cases such as dynamic baselining, alert correlation, event suppression, impact prediction, and automated context enrichment.
• Create automation, runbooks, and controlled self-healing mechanisms with appropriate safeguards, rollback strategies, and traceability.
• Implement AIOps remediation patterns that link observability signals to runbooks, tickets, ChatOps actions, and human-approved recovery procedures.
• Propel enhancements in system reliability, performance, scalability, and resilience through engineering-led initiatives.
• Collaborate with engineering teams to enhance deployment safety, operational readiness, and production stability.
• Maintain comprehensive runbooks, documentation, and knowledge bases to boost on-call efficiency and knowledge sharing.
• Assist in capacity planning, performance optimization, and SLO-based reliability practices.
• Enforce security, access controls, and operational guardrails across systems and automation.
• Oversee AIOps implementation from use-case definition to production deployment, encompassing telemetry readiness, model/rule configuration, integration testing, operational validation, adoption tracking, and continuous tuning.
• 6+ years of experience in SRE, AIOps, or production engineering within large-scale cloud environments.
• Strong proficiency in Linux/Unix, networking, distributed systems, and cloud platforms (AWS/Azure/GCP).
• Expertise in ELK/OpenSearch, including: log ingestion (Logstash / Beats), Elasticsearch index design, scaling, and tuning; advanced Kibana querying and debugging; dashboards, alerts, and observability patterns for production systems.
• Practical experience with logs, metrics, and tracing.
• Capability to prepare telemetry for AIOps implementation, including tagging, normalization, correlation keys, service mapping, and high-quality event metadata.
• Robust understanding of incident management, RCA, SLOs, and operational best practices.
• Good grasp of AIOps concepts: anomaly detection, alert correlation, and intelligent alerting.
• Hands-on experience in implementing AIOps workflows, including event correlation, signal enrichment, noise suppression, automated incident summaries, and governed remediation.
• Ability to integrate AIOps across observability tools, ITSM/ticketing systems, ChatOps, CMDB/runbook repositories, and automation platforms.
• Experience in measuring AIOps effectiveness through operational KPIs such as alert-noise reduction, quicker MTTD/MTTR, RCA quality, automation adoption, repeat usage, and linkage to business impact.
• Familiarity with Infrastructure as Code tools such as Terraform, Ansible, or similar.
• Preferred certifications: AWS DevOps Engineer, AWS ML Specialty, Google Cloud DevOps Engineer, Azure DevOps Engineer, Kubernetes/CKA, or relevant AI/ML, AIOps, observability, or cloud automation certifications.
• Employee Resource Groups to promote diverse voices.
• Coffee with Mark sessions – Our employees have the opportunity to engage with our CEO on significant and sometimes challenging topics ranging from mental health to work-life balance and current events.
• Microsoft Teams communities centered on wellness, art, pets, family, parenting, and more.
• Occasional special guests to discuss issues that affect our employee community.
DATAGROUP
Ambush
DuoKey
TEKsystems
Get handpicked remote jobs straight to your inbox weekly.