
Site Reliability Engineer
Posted 1 day ago

Posted 1 day ago
This is a fully remote position, open to applicants in India.
• Ensure the production platform remains reliable, scalable, and highly available.
• Establish and monitor SLIs, SLOs, and error budgets for essential services.
• Develop and enhance monitoring, alerting, logging, and tracing capabilities utilizing Prometheus, Grafana, Loki, Jaeger, and OpenTelemetry.
• Oversee and optimize Kubernetes clusters and containerized workloads.
• Automate routine operational tasks using Python, Go, or Bash.
• Construct and uphold infrastructure as code through Terraform and Helm.
• Maintain and enhance CI/CD pipelines to facilitate safe and frequent deployments, including blue/green and canary rollouts.
• Engage in an on-call rotation, lead incident response efforts, troubleshoot production problems, and conduct blameless post-mortems.
• Monitor and optimize PostgreSQL performance, backups, and high availability in collaboration with the development team.
• Plan for capacity, assess resilience, and contribute to reducing cloud expenses.
• Work closely with development, DevOps, and QA teams located in Canada.
• 6–8 years of experience in SRE, DevOps, or production infrastructure roles.
• Extensive hands-on experience with Kubernetes and Docker in production settings.
• Familiarity with at least one major cloud platform (AWS, Azure, or GCP).
• Proficient with monitoring and observability tools (Prometheus, Grafana, ELK/Loki, Jaeger, or similar).
• Experience in infrastructure as code using Terraform, Helm, or equivalent tools.
• Scripting or coding proficiency in Python, Go, or Bash.
• Strong foundational knowledge of Linux and networking (DNS, load balancing, TCP/IP, TLS).
• Experience with CI/CD tools such as GitHub Actions, GitLab CI, Jenkins, or Azure DevOps.
• A solid understanding of microservices and distributed systems.
• Experience in managing production incidents and producing root-cause analyses.
• Knowledge of service mesh technologies (Istio, Consul, or Linkerd) — nice to have.
• PostgreSQL administration and performance tuning experience — nice to have.
• Familiarity with .NET Core applications in a production environment — nice to have.
• Experience with messaging systems such as Kafka or RabbitMQ — nice to have.
• Relevant certifications like CKA, CKAD, or AWS/Azure DevOps Engineer — nice to have.
• Experience in logistics, transportation, or another 24/7 operations setting — a plus.
• Competitive Salary
• Healthcare Benefit Package
• Career Growth
Karat
Career TEAM
Jobsity
Career TEAM
Get handpicked remote jobs straight to your inbox weekly.