
Site Reliability Engineer II
Posted 1 day ago

Posted 1 day ago
This is a fully remote position, open to applicants in India.
• Work collaboratively with development teams to design and implement scalable infrastructure solutions.
• Create and implement automated deployment and testing pipelines.
• Develop and sustain monitoring and alerting systems.
• Troubleshoot production incidents and escalate as necessary to minimize downtime and enhance reliability.
• Continuously enhance infrastructure and processes for better scalability and efficiency.
• Participate in and take ownership of on-call rotations to deliver 24/7 application support.
• Conduct routine maintenance and system upgrades.
• Strengthen security measures and ensure compliance with industry standards.
• Convey technical concepts to both technical and non-technical stakeholders effectively.
• Mentor and guide junior engineers.
• Stay updated with advancements in site reliability engineering and share insights.
• Identify opportunities for organizational improvements and suggest alternatives to optimize team structure and execution.
• Bachelor’s degree in Computer Engineering, Computer Science, or a related discipline.
• Over 5 years of experience in a similar capacity, ideally in a high-traffic, high-availability setting.
• Proficient in at least one programming language, such as Python, Ruby, Java, or Go.
• Strong knowledge of cloud infrastructure and related technologies, including AWS, GCP, Azure, Kubernetes, and Docker.
• Exceptional troubleshooting and problem-solving abilities.
• Experience with automation and configuration management tools like Chef, Ansible, Puppet, or Terraform.
• Familiarity with monitoring and alerting tools such as Prometheus, Grafana, or Nagios.
• Strong communication and interpersonal skills.
• Ability to manage ambiguity, set clear expectations, and excel in a fast-paced, dynamic environment.
• Solid understanding of computer science fundamentals related to distributed systems and networks.
• Experience or familiarity with managing, tuning, and optimizing databases and queries.
• Previous experience in a leadership or senior-level site reliability engineering position.
• Familiarity with backend technologies and frameworks.
• Proven experience driving technical decisions and implementing organizational changes.
• An inclusive and diverse work environment.
• A remote work setting.
• Competitive salary.
• Potential share options for specific roles.
• Regular training opportunities.
• Annual learning stipend.
• A high degree of autonomy.
• Mentorship opportunities.
• Ambitious goals that support both personal and company growth.
Sprezzatura
Clinician Nexus
Lenovo
CEQUENS
Get handpicked remote jobs straight to your inbox weekly.