
Site Reliability Engineer
Posted Jul 19

Posted Jul 19
This is a fully remote position, open to applicants in Canada.
• Design, manage, and enhance dependable infrastructure for AI training and inference tasks.
• Take ownership of and automate operational processes in key areas such as networking, compute resources, storage, GPU/server setup, or AI platforms.
• Develop monitoring systems, alerts, runbooks, and incident-response strategies to simplify system operations.
• Identify and troubleshoot performance, capacity, and reliability challenges across hardware, operating systems, networks, schedulers, and distributed workloads.
• Collaborate closely with ML, research, and platform teams to convert workload requirements into actionable infrastructure enhancements.
• Enhance provisioning, configuration management, testing, and deployment automation processes.
• Assist in planning cluster expansion, capacity distribution, upgrades, and lifecycle management.
• Foster a culture of reliability through thorough documentation, post-incident analyses, and practical engineering standards.
• Minimum of 4 years of experience in site reliability engineering, infrastructure engineering, systems engineering, or a comparable production operations role.
• Proven hands-on proficiency in at least one of the following areas:
• Networking, encompassing firewalls, switching, routing, ASN/BGP configuration, or InfiniBand.
• Cluster and systems management using Kubernetes, SLURM, MAAS, or similar platforms.
• Distributed storage, specifically Ceph.
• GPU and server management, including CUDA drivers, firmware, BIOS, and hardware diagnostics.
• Infrastructure for AI training or model-serving.
• Experience managing production systems with an emphasis on availability, performance, security, and automation.
• Solid Linux administration and scripting capabilities.
• A methodical approach to troubleshooting across various layers of a complex system.
• Strong written and verbal communication skills, with the ability to collaborate effectively within a distributed team.
• Competitive salary and performance-based bonuses.
• Comprehensive health benefits, including medical, dental, and vision coverage.
• Flexible working hours and remote work options.
• Opportunities for professional development and continuous learning.
• A collaborative and inclusive work environment.
The Codest
IRIUM
Sólides
Resilinc
Get handpicked remote jobs straight to your inbox weekly.