
Site Reliability Engineer
Posted 3 hours ago

Posted 3 hours ago
This is a fully remote position, open to applicants in India.
• Participate in an on-call rotation addressing production availability incidents and assisting service engineers with customer-related issues.
• Utilize on-call shifts to mitigate the recurrence of incidents.
• Manage infrastructure using Ansible, Puppet, Terraform, and Kubernetes.
• Set up monitoring and alerting systems to identify symptoms rather than full outages.
• Keep a detailed record of all actions taken, ensuring that insights become repeatable practices and automation.
• Enhance the deployment process.
• Design, construct, and sustain core infrastructure capable of scaling to hundreds of thousands of concurrent users.
• Troubleshoot production issues across various services and stack levels.
• Strategize infrastructure expansion.
• Code infrastructure automation utilizing Ansible and Terraform.
• Enhance Prometheus monitoring or create new metrics.
• Assist release managers in deploying and resolving issues with new application software versions.
• Plan and implement the migration from AWS virtual machines to cloud-native, containerized deployments on Kubernetes (EKS).
• Build relationships with product teams and establish their SRE KPIs.
• Embrace a cloud-first mindset, regardless of the public cloud provider.
• Prioritize security in all considerations.
• Have a comprehensive understanding of systems, including edge cases, failure modes, behaviors, and specific implementations.
• Proficient in both Linux and Windows operating systems.
• Familiar with configuration management systems such as Ansible or Puppet.
• Strong programming capabilities in Python, Java, Golang, or Node.js.
• Ability to collaborate and communicate in an asynchronous manner.
• Thoroughly document all work.
• Possess a proactive attitude and a willingness to resolve issues in dysfunctional systems.
• Experience with Nginx, HAProxy, Docker, Kubernetes, Terraform, or similar technologies.
• Competitive salary and performance-based bonuses.
• Comprehensive health, dental, and vision insurance.
• Flexible working hours and remote work options.
• Opportunities for professional development and continuous learning.
• Collaborative and inclusive work environment.
Endava
Jones Lang LaSalle Americas, Inc.
NVIDIA
Entarian
Get handpicked remote jobs straight to your inbox weekly.