Remotery

Senior Site Reliability Engineer, SRE

Posted 2 days ago

This is a fully remote position, open to applicants in Bulgaria.

📋 Description

• Collaborate with international teams across various locations to tackle technical challenges and enhance processes.

• Design, implement, manage, and troubleshoot cloud and AI infrastructure solutions utilizing open source software.

• Deploy AI infrastructure based on NVIDIA-certified hardware, following engineering architecture and implementation designs.

• Work alongside stakeholders to collect and refine technical specifications.

• Enhance system performance, reliability, and scalability.

• Diagnose, debug, and resolve intricate technical issues.

• Engage in code reviews to uphold high quality standards.

• Keep abreast of trends and best practices in cloud operations and development.

• Create and implement AI-driven automation throughout the DevOps lifecycle, from code development to maintenance.

• Facilitate knowledge transfer to clients during the delivery phases.

• Mentor team members as well as Mirantis customers.

• Collaborate with stakeholders to define technical strategies and ensure the smooth integration of cloud and software services.


⛳️ Requirements

• Over 5 years of professional experience in DevOps, focusing on cloud and infrastructure technologies.

• Proficient in Kubernetes and/or OpenStack.

• Experience with high-performance data center processing, networking, and storage solutions.

• Familiarity with Golang, along with working knowledge of Python and JavaScript.

• In-depth understanding of distributed systems, microservices architecture, and CI/CD pipelines.

• Strong problem-solving and debugging capabilities across networking, storage, Linux, and Kubernetes.

• Knowledge of performance optimization and security measures.

• Capacity to lead technical tasks and collaborate with diverse teams.

• Ability to make independent judgment calls while engaging directly with customers.

• Exceptional written and verbal English communication skills.

• Excellent customer-facing communication abilities.

• Commitment to innovation, ongoing learning, and delivering high-quality results.

• Willingness to travel up to 25%, including internationally.

• Bachelor’s degree in Computer Science or a related field, or equivalent experience.

• Minimum of 5 years of experience in DevOps or Software Development, or similar roles.

• Nice to have: experience in network and/or storage architecture.

• Nice to have: experience with high-performance computing or GPU infrastructure, including GPU scheduling, MIG/vGPU, RDMA/RoCE or InfiniBand, NVLink, DCGM health-checking, GPU driver/firmware lifecycle, or NVIDIA AI Enterprise.

• Nice to have: presence in the open source community, upstream contributions, or conference presentations.

• Nice to have: familiarity with Rancher, OpenShift, or VMware.


🏝️ Benefits

• Opportunities for professional development and training.

• Attend conferences and participate in working groups.

• Company outings, happy hours, hackathons, and tech talks.

• Competitive compensation package with a comprehensive benefits plan.

• Chance to collaborate with passionate and talented colleagues.

• Open-source innovation environment.

• Dynamic workplace valuing openness, collaboration, risk-taking, and continuous growth.

People also viewed

CWILL14 hours ago

DevOps/SRE Engineer, Bilingual Mandarin

US flagCalifornia, +4 more statesFull-timeDevOps & Site Reliability Engineer (SRE)$100k – $130k/year
ApplyView job
a3715 hours ago

Forward Deployed DevOps Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
GT16 hours ago

Site Reliability Engineer, SRE

PL flagPoland, +2 more statesFreelanceDevOps & Site Reliability Engineer (SRE)
ApplyView job
Sigma Software Group16 hours ago

DevOps Engineer

PL flagPoland OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Applaudo16 hours ago

Google Cloud DevOps Engineer – Temporary Contract

CO flagColombia OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Branch16 hours ago

Cloud Operations Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$135k – $150k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers