
Site Reliability Engineer – II
Posted Jul 18

Posted Jul 18
This is a fully remote position, open to applicants in India.
• Collaborate with development teams to design and create scalable infrastructure.
• Work alongside development teams to establish automated deployment and testing pipelines.
• Develop and maintain monitoring and alerting systems to proactively identify and address potential issues.
• Troubleshoot and escalate production incidents to reduce downtime and enhance system reliability.
• Continuously enhance our infrastructure and processes to maximize scalability and efficiency.
• Participate in and take responsibility for on-call rotations as necessary to ensure 24/7 support for our applications.
• Conduct routine maintenance and upgrades as required to keep our systems current.
• Contribute to ongoing initiatives aimed at improving our security posture and ensuring compliance with industry standards.
• Effectively communicate complex technical concepts to both technical and non-technical stakeholders to facilitate informed decision-making.
• Mentor and guide junior engineers, promoting their professional development and enabling them to produce high-quality work.
• Stay informed about the latest advancements and trends in site reliability engineering, sharing knowledge and insights with the team.
• Identify opportunities for organizational improvements and suggest alternatives to optimize team structures and execution.
• Bachelor’s degree in Computer Engineering, Computer Science, or a related field.
• Over 5 years of experience in a similar position, ideally within a high-traffic, high-availability environment.
• Proficient in at least one programming language (Python, Ruby, Java, Go, etc.).
• Strong understanding of cloud infrastructure and associated technologies (AWS, GCP, Azure, Kubernetes, Docker, etc.).
• Excellent troubleshooting and problem-solving abilities.
• Experience with one or more automation and configuration management tools (Chef, Ansible, Puppet, Terraform, etc.).
• Familiarity with monitoring and alerting tools (Prometheus, Grafana, Nagios, etc.).
• Strong communication and interpersonal skills, facilitating effective collaboration with cross-functional teams.
• Ability to navigate ambiguity, establish clear expectations, and thrive in a fast-paced, dynamic environment.
• A solid understanding of computer science fundamentals related to distributed systems and networks.
• Inclusive and Diverse Environment: We cultivate an inclusive and diverse workplace that values innovation and provides remote opportunities.
• Competitive Compensation: Our compensation packages are highly competitive and may include potential share options for certain roles.
• Personal Growth and Development: We are dedicated to your personal and professional development, offering regular training and an annual learning stipend to help advance your career in a dynamic setting.
• Autonomy and Mentorship: You will enjoy significant autonomy in your role, backed by mentorship and ambitious goals that contribute to both your success and the company’s growth.
The Codest
IRIUM
Sólides
Resilinc
Get handpicked remote jobs straight to your inbox weekly.