
Senior Site Reliability Engineer, Kubernetes
Posted Jul 31

Posted Jul 31
This is a fully remote position, open to applicants in Canada.
• Participate in the design, development, and management of advanced cloud-based AI solutions utilizing the CNCF ecosystem, including Kubernetes.
• Primarily concentrate on deploying AI infrastructure using NVIDIA-certified hardware, adhering to the architectural and implementation designs developed by our engineering team.
• Guarantee the reliability, security, and performance of container infrastructure.
• Guide team members and Mirantis customers to provide high-quality software and services.
• Collaborate closely with stakeholders to establish technical strategies.
• Tackle complex challenges.
• Ensure the smooth integration of cloud and software services.
• Develop, implement, maintain, and troubleshoot cloud and AI infrastructure solutions based on open-source software.
• Work with stakeholders to collect and refine technical requirements.
• Enhance system performance, reliability, and scalability.
• Diagnose, debug, and resolve intricate technical issues.
• Engage in code reviews to uphold high-quality standards.
• Stay informed about industry trends and best practices in cloud operations and development.
• Design and implement AI-driven automation throughout the DevOps lifecycle, including code development and maintenance.
• Facilitate knowledge transfer to customers during delivery phases.
• Over 5 years of professional experience in DevOps, with a significant emphasis on Cloud, infrastructure technologies, and Kubernetes.
• Experience in high-performance data center processing, networking, and storage.
• Familiarity with Golang and a working knowledge of other programming languages (Python, JavaScript).
• Strong understanding of distributed systems, microservices architecture, and CI/CD pipelines.
• Exceptional problem-solving and debugging capabilities across networking and storage (hardware and software), Linux, and Kubernetes, with a focus on performance optimization and security.
• Proven ability to lead technical tasks and collaborate effectively with diverse teams.
• Comfortable making independent judgment calls when interacting directly with customers, often with limited day-to-day oversight.
• Excellent written and spoken English skills.
• Strong customer-facing communication abilities.
• A commitment to innovation, continuous learning, and delivering high-quality outcomes.
• Willingness to travel up to 25% if required, including internationally.
• Extensive experience in network and/or storage architecture (Nice to have).
• Experience with high-performance computing or GPU infrastructure (Nice to have).
• Practical experience with Openstack (Nice to have).
• Engagement in the open-source community, including upstream contributions and conference presentations (Nice to have).
• Prior experience with commercial container and virtual compute infrastructure platforms such as Rancher, Openshift, and VMware (Nice to have).
• Collaborate with an established Silicon Valley leader in the cloud infrastructure sector.
• Work alongside exceptionally passionate, talented, and engaging colleagues, assisting Fortune 500 and Global 2000 clients in implementing next-generation cloud technologies.
• Be part of cutting-edge, open-source innovation.
• Flourish in the dynamic environment of a young company where openness, collaboration, risk-taking, and continuous growth are valued.
• Opportunities for professional development and training.
• Attend conferences and working groups.
• Enjoy company outings, happy hours, hackathons, and tech talks.
• Receive a competitive compensation package accompanied by a robust benefits plan.
DATAGROUP
Ambush
DuoKey
TEKsystems
Get handpicked remote jobs straight to your inbox weekly.