Remotery

Senior Site Reliability Engineer, Kubernetes

Posted Jul 31

This is a fully remote position, open to applicants in Canada.

📋 Description

• Participate in the design, development, and management of advanced cloud-based AI solutions utilizing the CNCF ecosystem, including Kubernetes.

• Primarily concentrate on deploying AI infrastructure using NVIDIA-certified hardware, adhering to the architectural and implementation designs developed by our engineering team.

• Guarantee the reliability, security, and performance of container infrastructure.

• Guide team members and Mirantis customers to provide high-quality software and services.

• Collaborate closely with stakeholders to establish technical strategies.

• Tackle complex challenges.

• Ensure the smooth integration of cloud and software services.

• Develop, implement, maintain, and troubleshoot cloud and AI infrastructure solutions based on open-source software.

• Work with stakeholders to collect and refine technical requirements.

• Enhance system performance, reliability, and scalability.

• Diagnose, debug, and resolve intricate technical issues.

• Engage in code reviews to uphold high-quality standards.

• Stay informed about industry trends and best practices in cloud operations and development.

• Design and implement AI-driven automation throughout the DevOps lifecycle, including code development and maintenance.

• Facilitate knowledge transfer to customers during delivery phases.


⛳️ Requirements

• Over 5 years of professional experience in DevOps, with a significant emphasis on Cloud, infrastructure technologies, and Kubernetes.

• Experience in high-performance data center processing, networking, and storage.

• Familiarity with Golang and a working knowledge of other programming languages (Python, JavaScript).

• Strong understanding of distributed systems, microservices architecture, and CI/CD pipelines.

• Exceptional problem-solving and debugging capabilities across networking and storage (hardware and software), Linux, and Kubernetes, with a focus on performance optimization and security.

• Proven ability to lead technical tasks and collaborate effectively with diverse teams.

• Comfortable making independent judgment calls when interacting directly with customers, often with limited day-to-day oversight.

• Excellent written and spoken English skills.

• Strong customer-facing communication abilities.

• A commitment to innovation, continuous learning, and delivering high-quality outcomes.

• Willingness to travel up to 25% if required, including internationally.

• Extensive experience in network and/or storage architecture (Nice to have).

• Experience with high-performance computing or GPU infrastructure (Nice to have).

• Practical experience with Openstack (Nice to have).

• Engagement in the open-source community, including upstream contributions and conference presentations (Nice to have).

• Prior experience with commercial container and virtual compute infrastructure platforms such as Rancher, Openshift, and VMware (Nice to have).


🏝️ Benefits

• Collaborate with an established Silicon Valley leader in the cloud infrastructure sector.

• Work alongside exceptionally passionate, talented, and engaging colleagues, assisting Fortune 500 and Global 2000 clients in implementing next-generation cloud technologies.

• Be part of cutting-edge, open-source innovation.

• Flourish in the dynamic environment of a young company where openness, collaboration, risk-taking, and continuous growth are valued.

• Opportunities for professional development and training.

• Attend conferences and working groups.

• Enjoy company outings, happy hours, hackathons, and tech talks.

• Receive a competitive compensation package accompanied by a robust benefits plan.

People also viewed

DATAGROUP2 days ago

DevOps Engineer

DE flagGermany OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Ambush2 days ago

DevOps Engineer

BR flagBrazil OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
DuoKey2 days ago

DevOps Engineer

MU flagMauritius OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
TEKsystems3 days ago

SRE – CloudOps, Practice Architect II

US flagIllinois OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$148.2k – $222.4k/year
ApplyView job
TEKsystems3 days ago

SRE CloudOps Practice Architect II

US flagTexas OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$148.2k – $222.4k/year
ApplyView job
Level Data3 days ago

Senior DevOps Engineer

US flagMassachusetts OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$120k – $135k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers