Remotery

Site Reliability Engineer

atBosonRemoteCA flagCanadaFull-timeDevOps & Site Reliability Engineer (SRE)Mid-levelSenior$125k – $250k/year

Posted Jul 19

This is a fully remote position, open to applicants in Canada.

📋 Description

• Design, manage, and enhance dependable infrastructure for AI training and inference tasks.

• Take ownership of and automate operational processes in key areas such as networking, compute resources, storage, GPU/server setup, or AI platforms.

• Develop monitoring systems, alerts, runbooks, and incident-response strategies to simplify system operations.

• Identify and troubleshoot performance, capacity, and reliability challenges across hardware, operating systems, networks, schedulers, and distributed workloads.

• Collaborate closely with ML, research, and platform teams to convert workload requirements into actionable infrastructure enhancements.

• Enhance provisioning, configuration management, testing, and deployment automation processes.

• Assist in planning cluster expansion, capacity distribution, upgrades, and lifecycle management.

• Foster a culture of reliability through thorough documentation, post-incident analyses, and practical engineering standards.


⛳️ Requirements

• Minimum of 4 years of experience in site reliability engineering, infrastructure engineering, systems engineering, or a comparable production operations role.

• Proven hands-on proficiency in at least one of the following areas:

• Networking, encompassing firewalls, switching, routing, ASN/BGP configuration, or InfiniBand.

• Cluster and systems management using Kubernetes, SLURM, MAAS, or similar platforms.

• Distributed storage, specifically Ceph.

• GPU and server management, including CUDA drivers, firmware, BIOS, and hardware diagnostics.

• Infrastructure for AI training or model-serving.

• Experience managing production systems with an emphasis on availability, performance, security, and automation.

• Solid Linux administration and scripting capabilities.

• A methodical approach to troubleshooting across various layers of a complex system.

• Strong written and verbal communication skills, with the ability to collaborate effectively within a distributed team.


🏝️ Benefits

• Competitive salary and performance-based bonuses.

• Comprehensive health benefits, including medical, dental, and vision coverage.

• Flexible working hours and remote work options.

• Opportunities for professional development and continuous learning.

• A collaborative and inclusive work environment.

People also viewed

The Codest6 days ago

DevOps Engineer

PL flagPoland OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
IRIUM6 days ago

Ingeniero/a Cloud DevOps

ES flagSpain OnlyFull-timeDevOps & Site Reliability Engineer (SRE)€33k – €40k/year
ApplyView job
Sólides6 days ago

Senior DevOps Engineer

BR flagBrazil OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Resilinc6 days ago

Junior/Senior Site Reliability Engineer – Night Shift

IN flagIndia OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Verity Group6 days ago

Senior SRE / DevOps Engineer

Anywhere in the WorldFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
HOESSLER & HOESSLER6 days ago

DevOps Software Engineer – Career Ambitions

DE flagGermany OnlyFull-timeDevOps & Site Reliability Engineer (SRE)€65k – €75k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers