
Site Reliability Engineer
Posted 2 days ago

Posted 2 days ago
This is a fully remote position, open to applicants in India.
• Deliver on-call production assistance, execute rapid triage, manage escalations, and ensure service restoration.
• Diagnose production and customer-related issues, spearhead incident troubleshooting, and facilitate root-cause analysis.
• Utilize AIOps-supported RCA, log clustering, timeline reconstruction, and incident summarizations.
• Take ownership and enhance the ELK/OpenSearch observability framework.
• Design and uphold observability across logs, metrics, and traces.
• Develop workflows for alert notifications, anomaly detection, and incident enrichment.
• Implement dynamic baselining, alert correlation, event suppression, impact forecasting, and automated context enrichment.
• Create automation, runbooks, and controlled self-healing processes with safety measures, rollback strategies, and audit trails.
• Link observability signals to runbooks, tickets, ChatOps actions, and human-approved recovery procedures.
• Enhance system reliability, performance, scalability, and resilience.
• Collaborate with engineering teams to ensure deployment safety, operational readiness, and production stability.
• Maintain runbooks, documentation, and knowledge bases.
• Assist in capacity planning, performance optimization, and SLO-based reliability practices.
• Implement security measures, access controls, and operational guardrails.
• Oversee AIOps implementation from use-case definition to production rollout, validation, adoption tracking, and ongoing adjustments.
• Minimum of 4 years of experience in SRE, AIOps, or production engineering within large-scale cloud environments.
• Strong proficiency in Linux/Unix, networking, distributed systems, and AWS/Azure/GCP.
• In-depth knowledge of ELK/OpenSearch, including Logstash/Beats, Elasticsearch index design, scaling and tuning, Kibana querying and debugging, dashboards, and alerts.
• Hands-on experience with logs, metrics, and tracing methodologies.
• Capability to prepare telemetry for AIOps implementation, involving tagging, normalization, correlation keys, service mapping, and event metadata.
• Understanding of incident management, root-cause analysis, SLOs, and operational best practices.
• Practical experience with AIOps implementation, including anomaly detection, event correlation, signal enrichment, noise suppression, incident summaries, and governed remediation.
• Ability to implement integrations across observability tools, ITSM/ticketing systems, ChatOps, CMDB/runbook repositories, and automation platforms.
• Experience in assessing AIOps effectiveness using operational KPIs.
• Familiarity with Infrastructure as Code tools such as Terraform, Ansible, or similar.
• Preferred certifications include AWS DevOps Engineer, AWS ML Specialty, Google Cloud DevOps Engineer, Azure DevOps Engineer, Kubernetes/CKA, or relevant certifications in AI/ML, AIOps, observability, or cloud automation.
• Willingness to work shifts – EMEA shifts.
• Remote-first organization.
• Flexible remote work arrangement.
• Access to Employee Resource Groups.
• Participate in "Coffee with Mark" sessions.
• Engage in Microsoft Teams communities centered on wellness, art, pets, family, and parenting.
• Attend special guest discussions on topics affecting employees.
FourEnergy GmbH
ICF
Mastercam
C&S Informática
Get handpicked remote jobs straight to your inbox weekly.