
Site Reliability Engineer – Engineering Productivity
Posted Sep 2

Posted Sep 2
This is a fully remote position, open to applicants in India.
• Develop, securely and incrementally deploy, and manage essential production systems.
• Emphasize scalability, reliability, observability, performance, and security.
• Oversee, support, and improve developer experience across various services.
• Create automation to eliminate manual tasks and operate production systems effectively.
• Monitor and react to alerts, improve alerting systems, and implement automated alert handling.
• Generate and maintain incident-response runbooks.
• Construct and deploy new systems with scalability, reliability, and observability as key priorities.
• Diagnose platform and infrastructure challenges and assist Arista software engineers with issue resolution.
• Collaborate with third-party vendor support.
• Implement new systems in a phased approach.
• Compose postmortems and devise strategies to prevent future incidents.
• Plan and communicate maintenance windows for production systems.
• Identify infrastructure challenges that create workflow bottlenecks and constraints for product development teams.
• Design and apply solutions to address infrastructure challenges.
• Research and adopt best practices for infrastructure and platform management.
• Enhance fault tolerance and performance improvements to boost system availability.
• Analyze open-source system designs and implementation details to enhance triage and resolution processes.
• A minimum of a BSc in Computer Science or Engineering with 5 years of experience, an MS in Computer Science or Engineering with 5 years of experience, or equivalent professional experience.
• Proficiency in one or more languages such as Go, Python, or shell scripting to create medium-complexity automation workflows.
• Familiarity with Linux or UNIX from both administration and debugging perspectives.
• Practical experience in operating software systems at scale.
• Background in server provisioning, particularly from storage and networking viewpoints.
• Strong analytical and software troubleshooting abilities.
• Experience with infrastructure-as-code practices.
• Proficient in managing databases like MariaDB, PostgreSQL, or MongoDB.
• Hands-on experience with Docker and virtualization technologies such as KVM, QEMU, or Kata Containers.
• Experience managing monitoring stacks including Prometheus, Loki, Tempo, InfluxDB, Grafana, or Thanos.
• Experience overseeing Elasticsearch clusters.
• Experience managing Artifactory or Docker registries.
• Experience with CI/CD systems such as ArgoCD or Spinnaker.
• Familiarity with version control systems like Perforce or Gerrit.
• Experience with infrastructure-as-code frameworks such as Ansible.
• Experience in managing large Java applications.
• Background in storage infrastructure management, including NAS, SAN, or Ceph.
• Recognized as a Great Place to Work for engineering, diversity, compensation, and work-life balance.
• Engineers enjoy full ownership of their projects.
• A flat and efficient management structure.
• Opportunities to engage in various domains.
• Access to all areas of the company.
• An inclusive environment that values diverse thoughts and perspectives.
Sprezzatura
Clinician Nexus
Lenovo
MRSOOL | مرسول
Get handpicked remote jobs straight to your inbox weekly.