
Site Reliability Engineer – Engineering Productivity, DevOps
Posted Jul 23

Posted Jul 23
This is a fully remote position, open to applicants in Poland.
• Develop, deploy safely and incrementally, and manage critical production systems with an emphasis on scalability, reliability, observability, performance, and security.
• Oversee, support, and enhance the developer experience across various services.
• Create automation to eliminate repetitive tasks and efficiently manage production systems.
• Actively monitor, respond to, and refine alerts while establishing automated alert management.
• Generate and maintain incident response runbooks.
• Assess platform and infrastructure issues while assisting Arista software engineers in their troubleshooting efforts.
• Collaborate with third-party vendor support.
• Compose postmortem reports and develop solutions to prevent incident recurrence.
• Plan and communicate maintenance schedules for production systems.
• Collaborate with Arista’s product development teams to identify infrastructural challenges that hinder their workflows.
• Design and implement solutions to address these challenges.
• Research and adopt best practices in infrastructure and platform management to ensure secure, scalable, and fault-tolerant systems.
• Analyze the design and implementation specifics of open-source systems for improved troubleshooting and resolution.
• A minimum of a BSc in Computer Science or Engineering with 3 years of experience, an MS in Computer Science or Engineering with 3 years of experience, or equivalent professional experience.
• Proficiency in one or more programming languages such as Go, Python, or shell scripting to create medium complexity automation workflows.
• Familiarity with Linux (or UNIX) from both an administrative and debugging perspective.
• Practical experience in managing software systems (including infrastructure and complex applications) at scale.
• Experience in server provisioning, particularly from storage and networking viewpoints.
• Strong analytical and software troubleshooting abilities.
• Familiarity with infrastructure-as-code practices.
• Health insurance
• Retirement plans
• Paid time off
• Flexible work arrangements
• Professional development
CWILL
a37
GT
Sigma Software Group
Get handpicked remote jobs straight to your inbox weekly.