Senior Site Reliability Engineer, SRE

Posted 6 days ago

This is a fully remote position, open to applicants in Kazakhstan, +1 more country.

📋 Description

• Design, develop, and manage advanced cloud-based AI solutions utilizing the CNCF ecosystem, including Kubernetes.

• Implement AI infrastructure based on NVIDIA-certified hardware following engineering architecture and implementation blueprints.

• Maintain the reliability, security, and performance of container infrastructure.

• Mentor team members and Mirantis clients.

• Collaborate with internationally distributed teams to tackle technical challenges and enhance processes.

• Create, implement, sustain, and troubleshoot cloud and AI infrastructure solutions utilizing open source software.

• Work with stakeholders to collect and refine technical requirements.

• Enhance system performance, reliability, and scalability.

• Diagnose, debug, and resolve intricate technical issues.

• Participate in code review sessions.

• Stay abreast of industry trends and best practices in cloud operations and development.

• Design and deploy AI-driven automation throughout the DevOps lifecycle.

• Facilitate the transfer of knowledge to customers during delivery phases.

• Collaborate with stakeholders to establish technical strategies and address complex challenges.

• Ensure the seamless integration of cloud and software services.


⛳️ Requirements

• Over 5 years of professional experience in DevOps, emphasizing cloud and infrastructure technologies, including Kubernetes and/or OpenStack.

• Proficient in high-performance data center processing, networking, and storage.

• Familiarity with Golang and a working knowledge of other programming languages, such as Python and JavaScript.

• Strong understanding of distributed systems, microservices architecture, and CI/CD pipelines.

• Exceptional problem-solving and debugging skills across networking and storage, Linux, and Kubernetes.

• Knowledge in performance optimization and security.

• Capability to lead technical tasks and work effectively with diverse teams.

• Ability to make independent decisions when interacting with customers, often with minimal day-to-day supervision.

• Excellent proficiency in written and spoken English.

• Outstanding communication skills for customer interactions.

• Commitment to innovation, continuous learning, and delivering superior results.

• Willingness to travel up to 25% when necessary, including internationally.

• Bachelor's degree in Computer Science or a related field, or equivalent experience.

• Minimum of 5 years in DevOps or Software Development experience, or in a similar position.

• Nice to have: extensive experience in network and/or storage architecture.

• Nice to have: experience with high-performance computing or GPU infrastructure, including GPU scheduling, MIG/vGPU, RDMA/RoCE or InfiniBand fabrics, NVLink, DCGM health-checking, GPU driver/firmware lifecycle, or NVIDIA AI Enterprise.

• Nice to have: involvement in the open source community, including upstream contributions and conference presentations.

• Nice to have: familiarity with Rancher, OpenShift, and VMware.


🏝️ Benefits

• Opportunities for professional development and training.

• Participation in conferences and working groups.

• Company outings, happy hours, hackathons, and tech talks.

• Competitive compensation package accompanied by a robust benefits plan.

• Chance to collaborate with passionate, talented colleagues.

• Exposure to innovative, cutting-edge open-source technologies.

• High-energy environment that values openness, collaboration, risk-taking, and continuous growth.

People also viewed

Koniag Government Services2 days ago

Architect/DevOps Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
FP Markets (First Prudential Markets)3 days ago

Senior DevOps Engineer

AM flagArmenia, +4 more countriesFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Modern Campus3 days ago

Senior DevOps Engineer

US flagUnited States, +1 more countryFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
InRule3 days ago

Site Reliability Engineer

US flagUnited States, +1 more countryFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Thumbtack3 days ago

Senior Software Engineer, Site Reliability Engineering

US flagUnited States, +38 more locationsFull-timeDevOps & Site Reliability Engineer (SRE)$179.4k – $272.8k/year
ApplyView job
Thumbtack3 days ago

Senior Software Engineer, Site Reliability Engineering

CA flagCanada, +1 more countryFull-timeDevOps & Site Reliability Engineer (SRE)C$180.2k – C$233.2k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers