
Senior Site Reliability Engineer, SRE
Posted 6 days ago

Posted 6 days ago
This is a fully remote position, open to applicants in Kazakhstan, +1 more country.
• Design, develop, deploy, operate, maintain, and troubleshoot cloud and AI infrastructure solutions utilizing open source software.
• Implement AI infrastructure on NVIDIA-certified hardware in accordance with engineering architecture and implementation designs.
• Ensure the reliability, security, and performance of container infrastructure.
• Collaborate with internationally distributed teams to address technical challenges and enhance processes.
• Work closely with stakeholders to gather and refine technical requirements.
• Optimize system performance, reliability, and scalability.
• Diagnose, debug, and resolve complex technical issues.
• Engage in code reviews.
• Design and implement AI-driven automation throughout the DevOps lifecycle.
• Facilitate knowledge transfer to clients during delivery phases.
• Mentor team members and Mirantis customers.
• Define technical strategies and ensure the seamless integration of cloud and software services.
• Over 5 years of professional experience in DevOps, focusing on cloud and infrastructure technologies, including Kubernetes and/or OpenStack.
• Experience with high-performance data center processing, networking, and storage solutions.
• Familiarity with Golang and a working knowledge of Python and JavaScript.
• Strong understanding of distributed systems, microservices architecture, and CI/CD pipelines.
• Exceptional problem-solving and debugging capabilities across networking and storage, Linux, and Kubernetes.
• Knowledge of performance optimization and security measures.
• Ability to lead technical initiatives and collaborate with diverse teams.
• Capacity to make independent judgment calls while interacting directly with customers.
• Excellent proficiency in written and spoken English.
• Strong customer-facing communication skills.
• Commitment to innovation, continuous learning, and delivering top-notch results.
• Willingness to travel up to 25%, including internationally.
• Bachelor's degree in Computer Science or a related field, or equivalent experience.
• Minimum of 5 years in DevOps or Software Development roles, or similar positions.
• Preferred: experience in network and/or storage architecture.
• Preferred: background in high-performance computing or GPU infrastructure, including GPU scheduling, MIG/vGPU, RDMA/RoCE or InfiniBand, NVLink, DCGM health-checking, GPU driver/firmware lifecycle, or NVIDIA AI Enterprise.
• Preferred: active participation in the open source community, including upstream contributions and conference presentations.
• Preferred: familiarity with Rancher, OpenShift, and VMware.
• Opportunities for professional development and training.
• Participation in conferences and working groups.
• Company outings, happy hours, hackathons, and tech talks.
• Competitive salary package coupled with a robust benefits plan.
• Collaborate with exceptionally passionate, talented, and engaging colleagues.
• Work with Fortune 500 and Global 2000 clients on next-generation cloud technologies.
• Be part of cutting-edge, open-source innovation.
• Thrive in a high-energy environment that values openness, collaboration, risk-taking, and continuous growth.
Koniag Government Services
FP Markets (First Prudential Markets)
Modern Campus
InRule
Get handpicked remote jobs straight to your inbox weekly.