
Site Reliability Engineer
Posted Aug 18

Posted Aug 18
This is a fully remote position, open to applicants in Philippines.
• Take ownership of the availability and performance of production SaaS applications hosted on Azure across various geographic regions.
• Lead the troubleshooting and resolution of issues related to cloud infrastructure and applications, such as AKS failures, deployment rollbacks, ingress and networking challenges, and autoscaling difficulties.
• Engage in an on-call rotation, including weekends, and manage incident response from detection through to resolution.
• Enhance disaster recovery, failover, and incident management processes for multi-region deployments.
• Develop and maintain automation scripts and monitoring tools to minimize manual effort.
• Create post-incident reviews and root-cause analyses, and ensure preventive actions are completed.
• Collaborate with senior engineers and cross-functional teams to promote best practices in reliability, observability, and performance.
• Contribute to ongoing improvements in infrastructure, tools, and processes.
• Communicate effectively with customer-facing stakeholders during incidents through status updates and written incident summaries.
• 5+ years of pertinent experience in Site Reliability Engineering, DevOps, or Cloud Administration, with responsibility for production systems.
• Practical experience managing Azure environments, including AKS/Kubernetes, core Azure services, cloud networking, and fundamental cloud security.
• Strong understanding of monitoring, logging, and alerting methodologies, including tools like Datadog, Azure Monitor, or the ELK stack.
• Proficiency in troubleshooting by analyzing logs and stack traces using Datadog APM.
• Familiarity with firewalls, load balancers, VPNs, DNS, and routing protocols.
• Experience in automation and scripting with PowerShell, Python, or similar languages.
• Knowledge of cloud backup, redundancy, and disaster recovery strategies, including geo-redundant and multi-region setups.
• Accountability for the entire incident lifecycle, from detection to post-mortem analysis.
• Excellent written communication skills for delivering incident updates and summaries.
• Experience with AWS Cloud Platform is preferred.
• Familiarity with CI/CD tools such as Azure DevOps is preferred.
• Prior experience with infrastructure-as-code tools like Terraform or ARM templates is preferred.
• Previous experience operating SaaS products with regional tenant architectures is preferred.
• Successful candidates will undergo a thorough criminal background check, education verification, and employment verification.
• Competitive salaries
• Meaningful bonus program
• Healthcare insurance
• Pension/retirement matching
• Comprehensive life insurance
• Employee assistance program
• Time off plans
• Paid company holidays
• Career progression
• Culture of innovation
• Global, collaborative work environment
Ninja - نينجا
Dreamix
Scalingo
inDrive
Get handpicked remote jobs straight to your inbox weekly.