Site Reliability Engineer – Engineering Productivity

Posted Sep 2

This is a fully remote position, open to applicants in India.

📋 Description

• Develop, securely and incrementally deploy, and manage essential production systems.

• Emphasize scalability, reliability, observability, performance, and security.

• Oversee, support, and improve developer experience across various services.

• Create automation to eliminate manual tasks and operate production systems effectively.

• Monitor and react to alerts, improve alerting systems, and implement automated alert handling.

• Generate and maintain incident-response runbooks.

• Construct and deploy new systems with scalability, reliability, and observability as key priorities.

• Diagnose platform and infrastructure challenges and assist Arista software engineers with issue resolution.

• Collaborate with third-party vendor support.

• Implement new systems in a phased approach.

• Compose postmortems and devise strategies to prevent future incidents.

• Plan and communicate maintenance windows for production systems.

• Identify infrastructure challenges that create workflow bottlenecks and constraints for product development teams.

• Design and apply solutions to address infrastructure challenges.

• Research and adopt best practices for infrastructure and platform management.

• Enhance fault tolerance and performance improvements to boost system availability.

• Analyze open-source system designs and implementation details to enhance triage and resolution processes.


⛳️ Requirements

• A minimum of a BSc in Computer Science or Engineering with 5 years of experience, an MS in Computer Science or Engineering with 5 years of experience, or equivalent professional experience.

• Proficiency in one or more languages such as Go, Python, or shell scripting to create medium-complexity automation workflows.

• Familiarity with Linux or UNIX from both administration and debugging perspectives.

• Practical experience in operating software systems at scale.

• Background in server provisioning, particularly from storage and networking viewpoints.

• Strong analytical and software troubleshooting abilities.

• Experience with infrastructure-as-code practices.

• Proficient in managing databases like MariaDB, PostgreSQL, or MongoDB.

• Hands-on experience with Docker and virtualization technologies such as KVM, QEMU, or Kata Containers.

• Experience managing monitoring stacks including Prometheus, Loki, Tempo, InfluxDB, Grafana, or Thanos.

• Experience overseeing Elasticsearch clusters.

• Experience managing Artifactory or Docker registries.

• Experience with CI/CD systems such as ArgoCD or Spinnaker.

• Familiarity with version control systems like Perforce or Gerrit.

• Experience with infrastructure-as-code frameworks such as Ansible.

• Experience in managing large Java applications.

• Background in storage infrastructure management, including NAS, SAN, or Ceph.


🏝️ Benefits

• Recognized as a Great Place to Work for engineering, diversity, compensation, and work-life balance.

• Engineers enjoy full ownership of their projects.

• A flat and efficient management structure.

• Opportunities to engage in various domains.

• Access to all areas of the company.

• An inclusive environment that values diverse thoughts and perspectives.

People also viewed

Sprezzatura1 day ago

Salesforce Release Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$130k – $150k/year
ApplyView job
Clinician Nexus1 day ago

Manager, DevOps

US flagArizona, +13 more statesFull-timeDevOps & Site Reliability Engineer (SRE)$133.1k – $221.9k/year
ApplyView job
Lenovo1 day ago

CI/CD Engineer

US flagNorth Carolina OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$146.2k – $224.1k/year
ApplyView job
MRSOOL | مرسول1 day ago

Site Reliability Engineer II

IN flagIndia OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
CEQUENS1 day ago

DevOps Engineer

EG flagEgypt OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
In All Media1 day ago

DevOps Engineer – Cloud

BR flagBrazil, +5 more countriesFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers