Remotery

Senior Site Reliability Engineer, SRE

Posted 2 days ago

This is a fully remote position, open to applicants in Spain.

📋 Description

• Assist in the design, development, and operation of cloud-based AI solutions leveraging the CNCF ecosystem, particularly Kubernetes.

• Implement AI infrastructure utilizing NVIDIA-certified hardware in accordance with engineering architecture and implementation specifications.

• Ensure the container infrastructure's reliability, security, and performance.

• Provide mentorship to team members and Mirantis clients.

• Collaborate with internationally distributed teams on technical challenges and process enhancements.

• Develop, deploy, maintain, and troubleshoot cloud and AI infrastructure utilizing open-source software.

• Work alongside stakeholders to collect and refine technical requirements.

• Enhance system performance, reliability, and scalability.

• Diagnose, debug, and resolve intricate technical issues.

• Engage in code reviews.

• Keep abreast of trends and best practices in cloud operations and development.

• Design and implement AI-driven automation throughout the DevOps lifecycle.

• Facilitate knowledge transfer to clients during the delivery phases.

• Define technical strategies and ensure smooth integration of cloud and software services.


⛳️ Requirements

• Over 5 years of professional experience in DevOps, emphasizing cloud and infrastructure technologies.

• Proficiency in Kubernetes and/or OpenStack.

• Experience with high-performance data center processing, networking, and storage solutions.

• Familiarity with Golang and competent in Python, JavaScript, or other programming languages.

• Strong understanding of distributed systems, microservices architecture, and CI/CD pipelines.

• Excellent problem-solving and debugging abilities across networking, storage, Linux, and Kubernetes.

• Knowledge of performance optimization and security practices.

• Capacity to lead technical initiatives and collaborate with a diverse range of teams.

• Ability to make autonomous decisions when interacting directly with clients.

• Outstanding written and verbal English communication skills.

• Excellent customer-facing communication capabilities.

• Dedication to innovation, continuous learning, and delivering high-quality results.

• Willingness to travel up to 25%, including internationally.

• Bachelor's degree in Computer Science or a related field, or equivalent experience.

• At least 5 years of experience in DevOps or Software Development, or a similar role.

• Nice-to-have: experience in network and/or storage architecture.

• Nice-to-have: background in high-performance computing or GPU infrastructure, including GPU scheduling, MIG/vGPU, RDMA/RoCE or InfiniBand, NVLink, DCGM, GPU driver/firmware lifecycle, or NVIDIA AI Enterprise.

• Nice-to-have: involvement in the open-source community, upstream contributions, or conference presentations.

• Nice-to-have: experience with Rancher, OpenShift, or VMware.


🏝️ Benefits

• Opportunities for professional development and training.

• Participation in conferences and working groups.

• Company outings, happy hours, hackathons, and tech talks.

• Competitive compensation package complemented by a robust benefits plan.

• Option for remote work.

People also viewed

CWILL15 hours ago

DevOps/SRE Engineer, Bilingual Mandarin

US flagCalifornia, +4 more statesFull-timeDevOps & Site Reliability Engineer (SRE)$100k – $130k/year
ApplyView job
a3716 hours ago

Forward Deployed DevOps Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
GT17 hours ago

Site Reliability Engineer, SRE

PL flagPoland, +2 more statesFreelanceDevOps & Site Reliability Engineer (SRE)
ApplyView job
Sigma Software Group17 hours ago

DevOps Engineer

PL flagPoland OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Applaudo17 hours ago

Google Cloud DevOps Engineer – Temporary Contract

CO flagColombia OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Branch17 hours ago

Cloud Operations Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$135k – $150k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers