Remotery

Senior Site Reliability Engineer – Build

Posted Jun 15

This is a fully remote position, open to applicants in Canada.

📋 Description

• Implement infrastructure as code at scale by designing, executing, and maintaining infrastructure-as-code patterns with Terraform and Kubernetes, facilitating seamless deployment and operation for engineers.

• Develop and sustain comprehensive systems for monitoring, logging, and alerting. Take the lead in incident response initiatives, conduct thorough post-mortems, and promote continuous enhancements in system reliability.

• Collaborate with our Security team to integrate security measures across all layers of Build infrastructure, ensuring compliance across more than 100 jurisdictions while minimizing friction for developers and customers.

• Focus on performance and cost optimization by consistently enhancing system efficiency, resource utilization, and cloud expenses. Provide recommendations that bolster both reliability and unit economics.

• Seek out and eliminate manual operational tasks, building tools and processes that empower teams to operate effectively without increasing headcount.

• Partner with platform teams to ensure that APIs, MCP, and CLI are robust and observable, providing infrastructure insights that influence platform development.


⛳️ Requirements

• Extensive senior-level SRE experience: Proven background in Site Reliability Engineering, DevOps Engineering, or SysOps roles, with experience in establishing and managing production systems at scale.

• Expertise in Kubernetes and AWS: In-depth, hands-on experience with Kubernetes in production environments, coupled with solid foundational knowledge of AWS compute, networking, storage, and managed services.

• Proficient in infrastructure-as-code: Skilled in Terraform or similar IaC tools, you define infrastructure through code rather than utilizing console interfaces.

• Experience with CI/CD and deployment automation: Practical experience in setting up and managing GitLab, GitHub Actions, Jenkins, or similar systems, with an understanding of deployment strategies, rollback mechanisms, and safety precautions.

• Strong scripting and systems knowledge: Proficient in bash scripting and comfortable troubleshooting system-level issues, interpreting logs, and grasping Linux kernel fundamentals.

• Excellent communication skills: Capable of articulating complex infrastructure decisions to both technical and non-technical stakeholders, with a knack for producing clear runbooks and documentation.

• Preferred qualifications: Familiarity with at least one backend programming language (Elixir, Python, Go, Java, Node.js, etc.), experience in consultancy settings, knowledge of container registry and artifact management (ECR, Docker Hub, etc.), depth in observability stacks (Datadog, Prometheus, ELK, Grafana, or similar), and experience in scaling multi-tenant platforms.


🏝️ Benefits

• Work from anywhere.

• Flexible paid time off.

• Flexible working hours (we operate asynchronously).

• 16 weeks of paid parental leave.

• Access to mental health support services.

• Stock options.

• Learning budget.

• Home office budget and IT equipment.

• Budget allocated for local in-person social events or co-working spaces.

People also viewed

Ontrac Solutions2 days ago

Site Reliability Engineer

PK flagPakistan OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
CyberSheath2 days ago

Cloud Operations Engineer

US flagVirginia OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$110k – $127k/year
ApplyView job
Ontrac Solutions2 days ago

Site Reliability Engineer

IN flagIndia OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
NVIDIA2 days ago

Service Reliability Engineer

US flagTexas OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$168k – $333.5k/year
ApplyView job
Nagarro2 days ago

Senior Site Reliability Engineer, AWS Cloud

RO flagRomania OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Capgemini2 days ago

Senior DevOps Engineer

UA flagUkraine OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers