
Senior Site Reliability Engineer, SRE
Posted 2 days ago

Posted 2 days ago
This is a fully remote position, open to applicants in Latvia.
• Collaborate with international teams spread across different locations to tackle technical challenges and enhance processes.
• Create, implement, manage, and troubleshoot cloud and AI infrastructure solutions utilizing open-source software.
• Work alongside stakeholders to collect and refine technical requirements.
• Improve system performance, reliability, and scalability.
• Diagnose, debug, and resolve intricate technical issues.
• Participate in code reviews to uphold high-quality standards.
• Keep up to date with industry trends and best practices in cloud operations and development.
• Design and execute AI-driven automation throughout the DevOps lifecycle, including code development and maintenance.
• Support knowledge transfer to clients during delivery phases.
• Deploy AI infrastructure on NVIDIA-certified hardware in accordance with engineering architecture and implementation designs.
• Ensure the reliability, security, and performance of container infrastructure.
• Mentor team members and Mirantis customers.
• Define technical strategies and guarantee the seamless integration of cloud and software services.
• Over 5 years of professional experience in DevOps, emphasizing cloud and infrastructure technologies.
• Proficient in Kubernetes and/or OpenStack.
• Experience with high-performance data center processing, networking, and storage.
• Familiarity with Golang and a working knowledge of Python, JavaScript, and other programming languages.
• Strong understanding of distributed systems, microservices architecture, and CI/CD pipelines.
• Outstanding problem-solving and debugging abilities across networking, storage, Linux, and Kubernetes.
• Knowledge of performance optimization and security measures.
• Capability to lead technical tasks and work collaboratively with diverse teams.
• Ability to make independent judgment calls when engaging directly with customers.
• Excellent written and verbal communication skills in English.
• Strong customer-facing communication abilities.
• Commitment to innovation, continuous learning, and delivering high-quality results.
• Willingness to travel up to 25%, including international travel.
• Bachelor's degree in Computer Science or a related field, or equivalent experience.
• At least 5 years of experience in DevOps or Software Development, or a similar role.
• Nice to have: experience in network and/or storage architecture.
• Nice to have: background in high-performance computing or GPU infrastructure, including GPU scheduling, MIG/vGPU, RDMA/RoCE or InfiniBand, NVLink, DCGM health-checking, GPU driver/firmware lifecycle, or NVIDIA AI Enterprise.
• Nice to have: involvement in the open-source community, upstream contributions, or conference presentations.
• Nice to have: experience with Rancher, OpenShift, and VMware.
• Opportunities for professional development and training.
• Participation in conferences and working groups.
• Company outings, happy hours, hackathons, and tech talks.
• Competitive compensation package with a comprehensive benefits plan.
• Chance to work alongside passionate, talented, and engaging colleagues.
• An open-source innovation environment.
• A collaborative, high-energy company culture that values openness, risk-taking, and continuous growth.
CWILL
a37
GT
Sigma Software Group
Get handpicked remote jobs straight to your inbox weekly.