
Lead Software Engineer, Cloud Site Reliability
Posted 3 days ago

Posted 3 days ago
This is a fully remote position, open to applicants in India.
• Oversee 24x7 NOC operations with required rotational shifts, ensuring system availability and compliance with SLAs
• Serve as Major Incident Manager for P1/P2 incidents, facilitating triage, coordinating war room efforts, and communicating with stakeholders
• Implement and improve observability practices across logs, metrics, and traces
• Utilize Datadog and Azure Monitor for effective monitoring and alerting
• Promote proactive monitoring, alert tuning, anomaly detection, and AIOps initiatives
• Manage Azure infrastructure and AKS clusters, including troubleshooting, scaling, and performance optimization
• Develop automation and self-healing workflows using Terraform, ARM, Helm, Power Automate, and scripting
• Collaborate with engineering teams to enhance reliability, streamline deployment pipelines, and optimize cloud-native architecture
• Create dashboards and reports utilizing Power BI and ServiceNow
• Conduct monthly business reviews and prepare leadership reporting
• Mentor team members and promote process standardization and operational excellence
• Ensure availability, reliability, performance, emergency response, and capacity planning for Icertis SaaS applications and associated services
• Carry out infrastructure and access provisioning, upgrades, deployments, and change management
• Drive architectural enhancements to improve scalability and optimize overall costs
• 7–12 years of experience in CloudOps / SRE / NOC settings (24x7 operations)
• Strong knowledge of Azure Infrastructure (VMs, Networking, Storage)
• Practical experience with Azure Kubernetes Service (AKS), Kubernetes, and Docker
• Extensive experience with monitoring and observability tools (Datadog, Azure Monitor)
• Proven track record in Incident Management / Major Incident Handling and monthly reporting
• Familiarity with Infrastructure as Code (Terraform, ARM templates, Helm)
• Proficient scripting skills in PowerShell, Python, or Bash
• Experience with ServiceNow (Incident, Problem, Change modules, and dashboards)
• Solid understanding of distributed systems and cloud-native architecture
• Exceptional communication, leadership, and problem-solving abilities
• Experience in multi-cloud environments (Azure/AWS)
• Familiarity with AIOps / predictive monitoring / self-healing systems
• Azure / Datadog / Kubernetes certifications are preferred
• Bachelor's Degree
• Opportunity to work in a dynamic and innovative environment
• Competitive salary and benefits package
• Professional development and career growth opportunities
• Collaborative team culture with a focus on excellence
FourEnergy GmbH
ICF
Mastercam
C&S Informática
Get handpicked remote jobs straight to your inbox weekly.