Remotery

Senior Site Reliability Engineer, SRE

Posted 2 days ago

This is a fully remote position, open to applicants in Latvia.

📋 Description

• Collaborate with international teams spread across different locations to tackle technical challenges and enhance processes.

• Create, implement, manage, and troubleshoot cloud and AI infrastructure solutions utilizing open-source software.

• Work alongside stakeholders to collect and refine technical requirements.

• Improve system performance, reliability, and scalability.

• Diagnose, debug, and resolve intricate technical issues.

• Participate in code reviews to uphold high-quality standards.

• Keep up to date with industry trends and best practices in cloud operations and development.

• Design and execute AI-driven automation throughout the DevOps lifecycle, including code development and maintenance.

• Support knowledge transfer to clients during delivery phases.

• Deploy AI infrastructure on NVIDIA-certified hardware in accordance with engineering architecture and implementation designs.

• Ensure the reliability, security, and performance of container infrastructure.

• Mentor team members and Mirantis customers.

• Define technical strategies and guarantee the seamless integration of cloud and software services.


⛳️ Requirements

• Over 5 years of professional experience in DevOps, emphasizing cloud and infrastructure technologies.

• Proficient in Kubernetes and/or OpenStack.

• Experience with high-performance data center processing, networking, and storage.

• Familiarity with Golang and a working knowledge of Python, JavaScript, and other programming languages.

• Strong understanding of distributed systems, microservices architecture, and CI/CD pipelines.

• Outstanding problem-solving and debugging abilities across networking, storage, Linux, and Kubernetes.

• Knowledge of performance optimization and security measures.

• Capability to lead technical tasks and work collaboratively with diverse teams.

• Ability to make independent judgment calls when engaging directly with customers.

• Excellent written and verbal communication skills in English.

• Strong customer-facing communication abilities.

• Commitment to innovation, continuous learning, and delivering high-quality results.

• Willingness to travel up to 25%, including international travel.

• Bachelor's degree in Computer Science or a related field, or equivalent experience.

• At least 5 years of experience in DevOps or Software Development, or a similar role.

• Nice to have: experience in network and/or storage architecture.

• Nice to have: background in high-performance computing or GPU infrastructure, including GPU scheduling, MIG/vGPU, RDMA/RoCE or InfiniBand, NVLink, DCGM health-checking, GPU driver/firmware lifecycle, or NVIDIA AI Enterprise.

• Nice to have: involvement in the open-source community, upstream contributions, or conference presentations.

• Nice to have: experience with Rancher, OpenShift, and VMware.


🏝️ Benefits

• Opportunities for professional development and training.

• Participation in conferences and working groups.

• Company outings, happy hours, hackathons, and tech talks.

• Competitive compensation package with a comprehensive benefits plan.

• Chance to work alongside passionate, talented, and engaging colleagues.

• An open-source innovation environment.

• A collaborative, high-energy company culture that values openness, risk-taking, and continuous growth.

People also viewed

CWILL14 hours ago

DevOps/SRE Engineer, Bilingual Mandarin

US flagCalifornia, +4 more statesFull-timeDevOps & Site Reliability Engineer (SRE)$100k – $130k/year
ApplyView job
a3715 hours ago

Forward Deployed DevOps Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
GT16 hours ago

Site Reliability Engineer, SRE

PL flagPoland, +2 more statesFreelanceDevOps & Site Reliability Engineer (SRE)
ApplyView job
Sigma Software Group16 hours ago

DevOps Engineer

PL flagPoland OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Applaudo16 hours ago

Google Cloud DevOps Engineer – Temporary Contract

CO flagColombia OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Branch16 hours ago

Cloud Operations Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$135k – $150k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers