Lead Software Engineer, Cloud Site Reliability

Posted 3 days ago

This is a fully remote position, open to applicants in India.

📋 Description

• Oversee 24x7 NOC operations with required rotational shifts, ensuring system availability and compliance with SLAs

• Serve as Major Incident Manager for P1/P2 incidents, facilitating triage, coordinating war room efforts, and communicating with stakeholders

• Implement and improve observability practices across logs, metrics, and traces

• Utilize Datadog and Azure Monitor for effective monitoring and alerting

• Promote proactive monitoring, alert tuning, anomaly detection, and AIOps initiatives

• Manage Azure infrastructure and AKS clusters, including troubleshooting, scaling, and performance optimization

• Develop automation and self-healing workflows using Terraform, ARM, Helm, Power Automate, and scripting

• Collaborate with engineering teams to enhance reliability, streamline deployment pipelines, and optimize cloud-native architecture

• Create dashboards and reports utilizing Power BI and ServiceNow

• Conduct monthly business reviews and prepare leadership reporting

• Mentor team members and promote process standardization and operational excellence

• Ensure availability, reliability, performance, emergency response, and capacity planning for Icertis SaaS applications and associated services

• Carry out infrastructure and access provisioning, upgrades, deployments, and change management

• Drive architectural enhancements to improve scalability and optimize overall costs


⛳️ Requirements

• 7–12 years of experience in CloudOps / SRE / NOC settings (24x7 operations)

• Strong knowledge of Azure Infrastructure (VMs, Networking, Storage)

• Practical experience with Azure Kubernetes Service (AKS), Kubernetes, and Docker

• Extensive experience with monitoring and observability tools (Datadog, Azure Monitor)

• Proven track record in Incident Management / Major Incident Handling and monthly reporting

• Familiarity with Infrastructure as Code (Terraform, ARM templates, Helm)

• Proficient scripting skills in PowerShell, Python, or Bash

• Experience with ServiceNow (Incident, Problem, Change modules, and dashboards)

• Solid understanding of distributed systems and cloud-native architecture

• Exceptional communication, leadership, and problem-solving abilities

• Experience in multi-cloud environments (Azure/AWS)

• Familiarity with AIOps / predictive monitoring / self-healing systems

• Azure / Datadog / Kubernetes certifications are preferred

• Bachelor's Degree


🏝️ Benefits

• Opportunity to work in a dynamic and innovative environment

• Competitive salary and benefits package

• Professional development and career growth opportunities

• Collaborative team culture with a focus on excellence

People also viewed

FourEnergy GmbH13 hours ago

Senior DevOps Engineer – Operations

DE flagGermany OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
ICF15 hours ago

Lead DevOps Engineer

US flagVirginia OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$131.3k – $223.1k/year
ApplyView job
Mastercam19 hours ago

DevSecOps Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
C&S Informática1 day ago

DevOps Engineer – Freelance/Contract, Mid-Level/Senior

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Convene1 day ago

Support and Deployment Engineer

SA flagSaudi Arabia OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Verity Group1 day ago

SRE Engineer

BR flagBrazil OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers