
Senior Site Reliability Engineer, SRE
Posted 2 days ago

Posted 2 days ago
This is a fully remote position, open to applicants in Czechia.
• Design, develop, implement, operate, maintain, and troubleshoot cloud-based AI infrastructure solutions utilizing open-source software.
• Implement AI infrastructure utilizing NVIDIA-certified hardware in accordance with engineering architecture and design specifications.
• Ensure the reliability, security, performance, and scalability of container-based infrastructure.
• Collaborate with international teams located in various regions to address technical challenges and enhance processes.
• Work with stakeholders to collect and refine technical requirements.
• Diagnose, debug, and resolve intricate technical issues related to networking, storage, Linux, and Kubernetes.
• Engage in code review processes.
• Design and execute AI-driven automation throughout the DevOps lifecycle.
• Facilitate knowledge transfer to clients during the delivery phases.
• Provide mentorship to team members and Mirantis customers.
• Establish technical strategies and make autonomous technical decisions while collaborating with customers.
• Keep abreast of trends and best practices in cloud operations and development.
• Over 5 years of professional experience in DevOps, concentrating on cloud and infrastructure technologies, including Kubernetes and/or OpenStack.
• Experience with high-performance data center processing, networking, and storage solutions.
• Familiarity with Golang along with working knowledge of Python, JavaScript, and other programming languages.
• Strong understanding of distributed systems, microservices architecture, and CI/CD pipelines.
• Excellent problem-solving and debugging capabilities in networking, storage, Linux, and Kubernetes.
• Knowledge of performance optimization and security best practices.
• Capacity to lead technical initiatives and collaborate with diverse teams.
• Ability to make independent judgment calls when engaging directly with customers with minimal daily supervision.
• Proficient written and spoken English communication skills.
• Outstanding customer-facing communication abilities.
• Dedication to innovation, continuous learning, and delivering high-quality outcomes.
• Willingness to travel up to 25%, including internationally if necessary.
• Bachelor's degree in Computer Science or a related discipline, or equivalent experience.
• At least 5 years of experience in DevOps or Software Development, or a similar role.
• Extensive experience in network and/or storage architecture is a plus.
• Experience with high-performance computing or GPU infrastructure is advantageous.
• Involvement in the open-source community is a plus.
• Familiarity with Rancher, OpenShift, and VMware is beneficial.
• Opportunities for professional development and training.
• Participation in conferences and working groups.
• Company outings, happy hours, hackathons, and tech talks.
• Competitive compensation package accompanied by a robust benefits plan.
• Collaborate with passionate, talented, and engaging colleagues.
• Be part of innovative open-source projects.
• An environment that values openness, collaboration, risk-taking, and continuous growth.
CWILL
a37
GT
Sigma Software Group
Get handpicked remote jobs straight to your inbox weekly.