
Senior Site Reliability Engineer, SRE
Posted 6 days ago

Posted 6 days ago
This is a fully remote position, open to applicants in Kazakhstan, +1 more country.
• Contribute to the design, development, and operation of advanced cloud-based AI solutions utilizing the CNCF ecosystem, including Kubernetes.
• Deploy AI infrastructure utilizing NVIDIA-certified hardware in accordance with engineering architecture and implementation designs.
• Ensure the reliability, security, and performance of container infrastructure.
• Mentor team members and Mirantis customers to provide high-quality software and services.
• Collaborate with international teams to address technical challenges and enhance processes.
• Develop, implement, maintain, and troubleshoot cloud and AI infrastructure solutions leveraging open-source software.
• Engage with stakeholders to gather and refine technical requirements.
• Optimize system performance, reliability, and scalability.
• Diagnose, debug, and resolve complex technical issues.
• Participate in code reviews to uphold high quality standards.
• Stay updated on industry trends and best practices in cloud operations and development.
• Design and implement AI-driven automation throughout the DevOps lifecycle, including code development and maintenance.
• Facilitate knowledge transfer to customers during delivery phases.
• Collaborate closely with stakeholders to define technical strategies, tackle complex challenges, and ensure seamless integration of cloud and software services.
• Over 5 years of professional experience in DevOps, with a strong emphasis on cloud and infrastructure technologies, including Kubernetes and/or OpenStack.
• Experience in high-performance data center processing, networking, and storage.
• Familiarity with Golang and working knowledge of other programming languages (Python, JavaScript).
• Strong understanding of distributed systems, microservices architecture, and CI/CD pipelines.
• Exceptional problem-solving and debugging capabilities across networking and storage (hardware and software), Linux, and Kubernetes, with a focus on performance optimization and security.
• Proven ability to lead technical tasks and collaborate effectively with diverse teams.
• Comfortable making independent judgment calls when interacting directly with customers, often with minimal day-to-day supervision.
• Excellent written and spoken English skills.
• Strong customer-facing communication abilities.
• Dedication to innovation, continuous learning, and delivering high-quality results.
• Willingness to travel up to 25% if necessary, including internationally.
• Bachelor's degree in Computer Science or a related field, or equivalent experience.
• Extensive experience in network and/or storage architecture.
• Experience with high-performance computing or GPU infrastructure, including GPU scheduling, MIG/vGPU, RDMA/RoCE or InfiniBand fabrics, NVLink, DCGM health-checking, GPU driver/firmware lifecycle, or NVIDIA AI Enterprise.
• Active participation in the open-source community, including upstream contributions and conference presentations.
• Previous experience with Rancher, Openshift, and VMware.
• Work with a well-established leader in the cloud infrastructure industry based in Silicon Valley.
• Collaborate with exceptionally passionate, talented, and engaging colleagues, assisting Fortune 500 and Global 2000 customers in implementing next-generation cloud technologies.
• Be part of cutting-edge, open-source innovation.
• Thrive in a dynamic environment of a young company where openness, collaboration, risk-taking, and continuous growth are highly valued.
• Engage in professional development and training opportunities.
• Attend conferences and working groups.
• Participate in company outings, happy hours, hackathons, and tech talks.
• Enjoy a competitive compensation package complemented by a robust benefits plan.
Koniag Government Services
FP Markets (First Prudential Markets)
Modern Campus
InRule
Get handpicked remote jobs straight to your inbox weekly.