Remotery

Site Reliability Engineer

Posted 5 hours ago

This is a fully remote position, open to applicants in Costa Rica.

📋 Description

• Contribute to the resilience, performance enhancement, and system architecture across Backcountry's platform.

• Lead the resolution of critical incidents, ensuring that fixes are systematically implemented through thorough postmortems.

• Utilize AI-assisted engineering tools (such as Claude Code, GitHub Copilot, and MCP-based agents) to investigate, automate, and deploy solutions across infrastructure and application repositories.

• Minimize manual effort by designing and implementing automation solutions.

• Collaborate with fellow Site Reliability Engineers, developers, and architects to assess and apply best practices for existing and upcoming workloads.

• Monitor system performance and capacity, proactively addressing issues before they arise.

• Work with engineering teams to develop, deploy, and support new features.

• Create and maintain observability tools (including metrics, logs, traces, and profiles) along with SLI/SLO instrumentation for Backcountry services.

• Engage in FinOps initiatives across GCP and AWS, focusing on capacity planning and committed-use discount strategies.

• Participate in the on-call support rotation as part of the SRE team.


⛳️ Requirements

• Over 3 years of experience in supporting containerized production services, ideally with Kubernetes.

• More than 3 years of experience with Infrastructure as Code tools (such as Terraform, AWS CDK, Ansible, etc.).

• 3+ years of cloud experience in managing environments within Google Cloud Platform and/or AWS (multi-cloud experience, particularly with Azure/Entra, is a plus).

• Proficient in diagnosing issues and deploying bug fixes directly within application code, ensuring service reliability and stability.

• Comfortable performing thorough investigations across both infrastructure and application/software git repositories to trace issues from start to finish.

• Skilled in using AI-assisted coding tools (like Claude Code, GitHub Copilot) and MCP-based agents to expedite investigation, code review, and automation tasks.

• Strong knowledge of scripting and programming languages such as Bash, Python, and TypeScript/Node.js.

• Experience managing Linux (any major distribution) in production environments.

• Solid understanding of internet application protocols (including DHCP, DNS, HTTPS, SSH, etc.).

• Familiarity with DevOps (CI/CD) and SRE practices (SLOs, SLIs) as they relate to daily operations.

• Practical experience with observability tools (such as Grafana, Prometheus, Loki, OpenSearch, or similar) and SLI/SLO instrumentation.

• Knowledge of GitOps and Kubernetes packaging tools (like ArgoCD, Helm, Kustomize).

• Actively monitor emerging technology trends and developments, assessing those worth integrating into engineering practices.

• Bachelor's degree in computer science or a related field, or equivalent professional experience.

• Proficient in English communication skills, both spoken and written, at an advanced level.


🏝️ Benefits

• Competitive Benefits: We provide a comprehensive benefits package that includes primarily remote work, private medical and life insurance, additional paid time off, monthly allowances and reimbursements, employee discounts, and opportunities for professional development.

People also viewed

NBCUniversal5 hours ago

Senior DevOps Engineer

US flagNew York OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$110k – $165k/year
ApplyView job
Sophos5 hours ago

Software Engineer – SRE

IN flagIndia OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
ClanX5 hours ago

Senior DevOps Engineer

IN flagIndia OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
General Dynamics Information Technology5 hours ago

Dev Sec Ops Developer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$97.8k – $132.3k/year
ApplyView job
OpenCV5 hours ago

Senior Site Reliability Engineer

DE flagGermany OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
VPS5 hours ago

DevSecOps Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$80k – $145k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers