
Senior Site Reliability Engineer – SRE
Posted 6 days ago

Posted 6 days ago
This is a fully remote position, open to applicants in Kazakhstan, +1 more country.
• Collaborate with internationally distributed teams to tackle technical challenges and enhance processes.
• Create, implement, manage, and troubleshoot cloud and AI infrastructure solutions utilizing open source software.
• Set up AI infrastructure using NVIDIA-certified hardware following engineering architecture and implementation designs.
• Contribute to the design, development, and management of cloud-based AI solutions on the CNCF ecosystem, including Kubernetes.
• Ensure the container infrastructure's reliability, security, and performance.
• Provide mentorship to team members and Mirantis customers.
• Work alongside stakeholders to collect and refine technical requirements.
• Enhance system performance, reliability, and scalability.
• Diagnose, debug, and resolve intricate technical issues.
• Engage in code reviews.
• Keep up-to-date with industry trends and best practices in cloud operations and development.
• Design and implement AI-driven automation throughout the DevOps lifecycle.
• Support knowledge transfer to customers during delivery phases.
• Define technical strategies and ensure the smooth integration of cloud and software services.
• Over 5 years of professional experience in DevOps, particularly focusing on cloud and infrastructure technologies, including Kubernetes and/or OpenStack.
• Proficient in high-performance data center processing, networking, and storage.
• Familiarity with Golang and a working knowledge of other programming languages such as Python and JavaScript.
• Strong understanding of distributed systems, microservices architecture, and CI/CD pipelines.
• Exceptional problem-solving and debugging abilities across networking and storage (both hardware and software), Linux, and Kubernetes, with a focus on performance optimization and security.
• Proven capability to lead technical tasks and collaborate effectively with diverse teams.
• Comfortable making independent decisions when engaging directly with customers, often with minimal day-to-day supervision.
• Excellent proficiency in written and spoken English.
• Outstanding communication skills in customer-facing situations.
• A commitment to innovation, continuous learning, and delivering high-quality outcomes.
• Ability to travel up to 25% if required, including internationally.
• Bachelor's degree in Computer Science or a related field, or equivalent experience.
• Minimum of 5 years of experience in DevOps or Software Development, or a similar role.
• Nice-to-have: extensive experience in network and/or storage architecture.
• Nice-to-have: experience with high-performance computing or GPU infrastructure, including GPU scheduling, MIG/vGPU, RDMA/RoCE or InfiniBand fabrics, NVLink, DCGM health-checking, GPU driver/firmware lifecycle, or NVIDIA AI Enterprise.
• Nice-to-have: involvement in the open source community, including upstream contributions and conference presentations.
• Nice-to-have: prior experience with Rancher, OpenShift, and VMware.
• Opportunities for professional development and training.
• Participation in conferences and working groups.
• Company outings, happy hours, hackathons, and tech talks.
• Competitive compensation package along with a robust benefits plan.
• Availability of remote work.
Koniag Government Services
FP Markets (First Prudential Markets)
Modern Campus
InRule
Get handpicked remote jobs straight to your inbox weekly.