
Senior Site Reliability Engineer, SRE
Posted 6 days ago

Posted 6 days ago
This is a fully remote position, open to applicants in Kazakhstan, +1 more country.
• Contribute to the design, development, and operation of cloud-based AI solutions utilizing the CNCF ecosystem, including Kubernetes.
• Deploy AI infrastructure based on NVIDIA-certified hardware in accordance with engineering architecture and implementation designs.
• Ensure the reliability, security, and performance of the container infrastructure.
• Mentor team members and assist Mirantis customers.
• Collaborate with geographically distributed international teams to tackle technical challenges and enhance processes.
• Develop, implement, maintain, and troubleshoot cloud and AI infrastructure solutions leveraging open source software.
• Work with stakeholders to gather and refine technical requirements.
• Optimize system performance, reliability, and scalability.
• Troubleshoot, debug, and resolve intricate technical issues.
• Participate in code reviews.
• Stay updated with trends and best practices in cloud operations and development.
• Design and implement AI-driven automation throughout the DevOps lifecycle.
• Facilitate knowledge transfer to customers during the delivery phases.
• Collaborate with stakeholders to define technical strategies and address complex challenges.
• Ensure seamless integration of cloud and software services.
• Over 5 years of professional experience in DevOps, emphasizing cloud and infrastructure technologies, including Kubernetes and/or OpenStack.
• Experience in high-performance data center processing, networking, and storage.
• Familiarity with Golang and a working knowledge of Python and JavaScript.
• Strong understanding of distributed systems, microservices architecture, and CI/CD pipelines.
• Exceptional problem-solving and debugging abilities across networking and storage, Linux, and Kubernetes, with a focus on performance optimization and security.
• Proven capability to lead technical tasks and collaborate effectively with diverse teams.
• Comfortable making independent judgments when directly interacting with customers, often with limited day-to-day supervision.
• Excellent proficiency in written and spoken English.
• Strong customer-facing communication skills.
• Commitment to innovation, continuous learning, and delivering high-quality results.
• Ability to travel up to 25% if necessary, including internationally.
• Bachelor's degree in Computer Science or a related field, or equivalent experience.
• Minimum of 5 years of experience in DevOps, Software Development, or a similar role.
• Nice to have: extensive experience in network and/or storage architecture.
• Nice to have: experience with high-performance computing or GPU infrastructure, including GPU scheduling, MIG/vGPU, RDMA/RoCE or InfiniBand fabrics, NVLink, DCGM health-checking, GPU driver/firmware lifecycle, or NVIDIA AI Enterprise.
• Nice to have: participation in the open source community, including upstream contributions and conference presentations.
• Nice to have: experience with Rancher, OpenShift, and VMware.
• Opportunities for professional development and training.
• Attendance at conferences and working groups.
• Company outings, happy hours, hackathons, and tech talks.
• Competitive compensation package alongside a robust benefits plan.
• Availability for remote work.
Koniag Government Services
FP Markets (First Prudential Markets)
Modern Campus
InRule
Get handpicked remote jobs straight to your inbox weekly.