Site Reliability Engineer

Posted Aug 10

This is a fully remote position, open to applicants in Madagascar.

📋 Description

• Address production availability incidents as part of an on-call rotation.

• Assist service engineers with customer-related incidents.

• Utilize on-call shifts to mitigate recurring incidents.

• Manage infrastructure using Ansible, Puppet, Terraform, and Kubernetes.

• Develop symptom-based monitoring and alerting systems.

• Document procedures and transform findings into repeatable actions and automation.

• Enhance deployment processes.

• Design, construct, and maintain infrastructure capable of scaling to hundreds of thousands of concurrent users.

• Troubleshoot production issues across various services and stack levels.

• Strategize for infrastructure growth.

• Program infrastructure automation utilizing Ansible and Terraform.

• Refine Prometheus monitoring and create new metrics.

• Assist release managers in deploying and troubleshooting application software versions.

• Plan and execute the migration from AWS virtual machines to Kubernetes-based cloud-native deployments on EKS.

• Cultivate relationships with product teams and establish SRE KPIs.


⛳️ Requirements

• Possess a cloud-first and security-first mindset.

• Apply systems thinking that considers edge cases, failure modes, behaviors, and implementations.

• Be familiar with both Linux and Windows operating systems.

• Have knowledge of configuration-management systems such as Ansible or Puppet.

• Exhibit strong programming capabilities in Python, Java, Golang, or Node.js.

• Demonstrate the ability to collaborate and communicate asynchronously, as well as document work effectively.

• Take a proactive approach to remedying broken systems.

• Have experience with technologies like Nginx, HAProxy, Docker, Kubernetes, Terraform, or similar tools.


🏝️ Benefits

• Opportunities to work with cutting-edge technology.

• Collaborative and innovative work environment.

• Flexible remote work options.

People also viewed

Horizon3.ai14 hours ago

Staff Site Reliability Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$199.8k – $270k/year
ApplyView job
CLOUD MANTA GmbH14 hours ago

Senior DevOps Engineer, Containers & Private Cloud

DE flagGermany OnlyFull-timeDevOps & Site Reliability Engineer (SRE)€70k – €80k/year
ApplyView job
Stefanini LATAM14 hours ago

Senior DevOps

AR flagArgentina OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Akamai Technologies14 hours ago

Principal Site Reliability Engineer – Lead

PL flagPoland OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
PingWind Inc. (SDVOSB)14 hours ago

DevSecOps Engineer

US flagAlabama, +1 more stateFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Ad Hoc LLC14 hours ago

Staff DevOps Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$130k – $150k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers