
Senior Site Reliability Engineer, SRE
Posted 2 days ago

Posted 2 days ago
This is a fully remote position, open to applicants in Spain.
• Assist in the design, development, and operation of cloud-based AI solutions leveraging the CNCF ecosystem, particularly Kubernetes.
• Implement AI infrastructure utilizing NVIDIA-certified hardware in accordance with engineering architecture and implementation specifications.
• Ensure the container infrastructure's reliability, security, and performance.
• Provide mentorship to team members and Mirantis clients.
• Collaborate with internationally distributed teams on technical challenges and process enhancements.
• Develop, deploy, maintain, and troubleshoot cloud and AI infrastructure utilizing open-source software.
• Work alongside stakeholders to collect and refine technical requirements.
• Enhance system performance, reliability, and scalability.
• Diagnose, debug, and resolve intricate technical issues.
• Engage in code reviews.
• Keep abreast of trends and best practices in cloud operations and development.
• Design and implement AI-driven automation throughout the DevOps lifecycle.
• Facilitate knowledge transfer to clients during the delivery phases.
• Define technical strategies and ensure smooth integration of cloud and software services.
• Over 5 years of professional experience in DevOps, emphasizing cloud and infrastructure technologies.
• Proficiency in Kubernetes and/or OpenStack.
• Experience with high-performance data center processing, networking, and storage solutions.
• Familiarity with Golang and competent in Python, JavaScript, or other programming languages.
• Strong understanding of distributed systems, microservices architecture, and CI/CD pipelines.
• Excellent problem-solving and debugging abilities across networking, storage, Linux, and Kubernetes.
• Knowledge of performance optimization and security practices.
• Capacity to lead technical initiatives and collaborate with a diverse range of teams.
• Ability to make autonomous decisions when interacting directly with clients.
• Outstanding written and verbal English communication skills.
• Excellent customer-facing communication capabilities.
• Dedication to innovation, continuous learning, and delivering high-quality results.
• Willingness to travel up to 25%, including internationally.
• Bachelor's degree in Computer Science or a related field, or equivalent experience.
• At least 5 years of experience in DevOps or Software Development, or a similar role.
• Nice-to-have: experience in network and/or storage architecture.
• Nice-to-have: background in high-performance computing or GPU infrastructure, including GPU scheduling, MIG/vGPU, RDMA/RoCE or InfiniBand, NVLink, DCGM, GPU driver/firmware lifecycle, or NVIDIA AI Enterprise.
• Nice-to-have: involvement in the open-source community, upstream contributions, or conference presentations.
• Nice-to-have: experience with Rancher, OpenShift, or VMware.
• Opportunities for professional development and training.
• Participation in conferences and working groups.
• Company outings, happy hours, hackathons, and tech talks.
• Competitive compensation package complemented by a robust benefits plan.
• Option for remote work.
CWILL
a37
GT
Sigma Software Group
Get handpicked remote jobs straight to your inbox weekly.