Site Reliability Engineer

Posted Aug 10

This is a fully remote position, open to applicants in Pakistan.

📋 Description

• Participate in an on-call rotation to address production availability incidents and assist service engineers with customer-related issues.

• Utilize on-call shifts to mitigate the recurrence of incidents.

• Operate infrastructure using tools such as Ansible, Puppet, Terraform, and Kubernetes.

• Set up monitoring and alerting systems to report symptoms rather than complete outages.

• Document every step taken to ensure findings can be replicated and automated.

• Enhance the deployment process.

• Design, construct, and maintain core infrastructure capable of scaling to support hundreds of thousands of concurrent users.

• Troubleshoot production issues across various services and stack levels.

• Strategize infrastructure growth.

• Code infrastructure automation utilizing Ansible and Terraform.

• Enhance Prometheus monitoring and develop new metrics as needed.

• Assist release managers in deploying and troubleshooting new versions of application software.

• Plan and carry out the migration from AWS virtual machines to cloud-native, container-based deployments on Kubernetes (EKS).

• Cultivate relationships with product teams and establish SRE KPIs.


⛳️ Requirements

• Adopt a cloud-first mentality, regardless of the public cloud provider.

• Prioritize security in all considerations.

• Possess a comprehensive understanding of systems, including edge cases, failure modes, behaviors, and specific implementations.

• Familiarity with both Linux and Windows operating systems.

• Knowledge of configuration management systems such as Ansible or Puppet.

• Proficient programming skills in Python, Java, Golang, or Node.js.

• Ability to collaborate and communicate asynchronously, with thorough documentation of work.

• A proactive attitude and readiness to address and resolve malfunctioning systems.

• Experience with technologies such as Nginx, HAProxy, Docker, Kubernetes, Terraform, or similar tools.


🏝️ Benefits

• Competitive salary and performance-based bonuses.

• Comprehensive health, dental, and vision insurance.

• Flexible working hours and remote work options.

• Opportunities for professional development and training.

• Collaborative and innovative work environment.

People also viewed

Horizon3.ai14 hours ago

Staff Site Reliability Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$199.8k – $270k/year
ApplyView job
CLOUD MANTA GmbH14 hours ago

Senior DevOps Engineer, Containers & Private Cloud

DE flagGermany OnlyFull-timeDevOps & Site Reliability Engineer (SRE)€70k – €80k/year
ApplyView job
Stefanini LATAM14 hours ago

Senior DevOps

AR flagArgentina OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Akamai Technologies14 hours ago

Principal Site Reliability Engineer – Lead

PL flagPoland OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
PingWind Inc. (SDVOSB)14 hours ago

DevSecOps Engineer

US flagAlabama, +1 more stateFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Ad Hoc LLC14 hours ago

Staff DevOps Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$130k – $150k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers