Site Reliability Engineer

Posted 1 day ago

This is a fully remote position, open to applicants in India.

📋 Description

• Ensure the production platform remains reliable, scalable, and highly available.

• Establish and monitor SLIs, SLOs, and error budgets for essential services.

• Develop and enhance monitoring, alerting, logging, and tracing capabilities utilizing Prometheus, Grafana, Loki, Jaeger, and OpenTelemetry.

• Oversee and optimize Kubernetes clusters and containerized workloads.

• Automate routine operational tasks using Python, Go, or Bash.

• Construct and uphold infrastructure as code through Terraform and Helm.

• Maintain and enhance CI/CD pipelines to facilitate safe and frequent deployments, including blue/green and canary rollouts.

• Engage in an on-call rotation, lead incident response efforts, troubleshoot production problems, and conduct blameless post-mortems.

• Monitor and optimize PostgreSQL performance, backups, and high availability in collaboration with the development team.

• Plan for capacity, assess resilience, and contribute to reducing cloud expenses.

• Work closely with development, DevOps, and QA teams located in Canada.


⛳️ Requirements

• 6–8 years of experience in SRE, DevOps, or production infrastructure roles.

• Extensive hands-on experience with Kubernetes and Docker in production settings.

• Familiarity with at least one major cloud platform (AWS, Azure, or GCP).

• Proficient with monitoring and observability tools (Prometheus, Grafana, ELK/Loki, Jaeger, or similar).

• Experience in infrastructure as code using Terraform, Helm, or equivalent tools.

• Scripting or coding proficiency in Python, Go, or Bash.

• Strong foundational knowledge of Linux and networking (DNS, load balancing, TCP/IP, TLS).

• Experience with CI/CD tools such as GitHub Actions, GitLab CI, Jenkins, or Azure DevOps.

• A solid understanding of microservices and distributed systems.

• Experience in managing production incidents and producing root-cause analyses.

• Knowledge of service mesh technologies (Istio, Consul, or Linkerd) — nice to have.

• PostgreSQL administration and performance tuning experience — nice to have.

• Familiarity with .NET Core applications in a production environment — nice to have.

• Experience with messaging systems such as Kafka or RabbitMQ — nice to have.

• Relevant certifications like CKA, CKAD, or AWS/Azure DevOps Engineer — nice to have.

• Experience in logistics, transportation, or another 24/7 operations setting — a plus.


🏝️ Benefits

• Competitive Salary

• Healthcare Benefit Package

• Career Growth

People also viewed

Karat4 hours ago

Senior DevOps Engineer

IN flagIndia OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Career TEAM17 hours ago

DevOps Engineer

PH flagPhilippines OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Jobsity17 hours ago

Senior Site Reliability Engineer, AWS Multi-region

AR flagArgentina, +2 more countriesFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Career TEAM17 hours ago

DevOps Engineer

PH flagPhilippines OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Dev.Pro1 day ago

Senior DevOps Engineer

BR flagBrazil, +2 more countriesFreelanceDevOps & Site Reliability Engineer (SRE)
ApplyView job
EVT1 day ago

Senior DevSecOps Engineer

BR flagBrazil OnlyFreelanceDevOps & Site Reliability Engineer (SRE)
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers