
Senior Site Reliability Engineer, SRE
Posted 2 days ago

Posted 2 days ago
This is a fully remote position, open to applicants in Bulgaria.
• Collaborate with international teams across various locations to tackle technical challenges and enhance processes.
• Design, implement, manage, and troubleshoot cloud and AI infrastructure solutions utilizing open source software.
• Deploy AI infrastructure based on NVIDIA-certified hardware, following engineering architecture and implementation designs.
• Work alongside stakeholders to collect and refine technical specifications.
• Enhance system performance, reliability, and scalability.
• Diagnose, debug, and resolve intricate technical issues.
• Engage in code reviews to uphold high quality standards.
• Keep abreast of trends and best practices in cloud operations and development.
• Create and implement AI-driven automation throughout the DevOps lifecycle, from code development to maintenance.
• Facilitate knowledge transfer to clients during the delivery phases.
• Mentor team members as well as Mirantis customers.
• Collaborate with stakeholders to define technical strategies and ensure the smooth integration of cloud and software services.
• Over 5 years of professional experience in DevOps, focusing on cloud and infrastructure technologies.
• Proficient in Kubernetes and/or OpenStack.
• Experience with high-performance data center processing, networking, and storage solutions.
• Familiarity with Golang, along with working knowledge of Python and JavaScript.
• In-depth understanding of distributed systems, microservices architecture, and CI/CD pipelines.
• Strong problem-solving and debugging capabilities across networking, storage, Linux, and Kubernetes.
• Knowledge of performance optimization and security measures.
• Capacity to lead technical tasks and collaborate with diverse teams.
• Ability to make independent judgment calls while engaging directly with customers.
• Exceptional written and verbal English communication skills.
• Excellent customer-facing communication abilities.
• Commitment to innovation, ongoing learning, and delivering high-quality results.
• Willingness to travel up to 25%, including internationally.
• Bachelor’s degree in Computer Science or a related field, or equivalent experience.
• Minimum of 5 years of experience in DevOps or Software Development, or similar roles.
• Nice to have: experience in network and/or storage architecture.
• Nice to have: experience with high-performance computing or GPU infrastructure, including GPU scheduling, MIG/vGPU, RDMA/RoCE or InfiniBand, NVLink, DCGM health-checking, GPU driver/firmware lifecycle, or NVIDIA AI Enterprise.
• Nice to have: presence in the open source community, upstream contributions, or conference presentations.
• Nice to have: familiarity with Rancher, OpenShift, or VMware.
• Opportunities for professional development and training.
• Attend conferences and participate in working groups.
• Company outings, happy hours, hackathons, and tech talks.
• Competitive compensation package with a comprehensive benefits plan.
• Chance to collaborate with passionate and talented colleagues.
• Open-source innovation environment.
• Dynamic workplace valuing openness, collaboration, risk-taking, and continuous growth.
CWILL
a37
GT
Sigma Software Group
Get handpicked remote jobs straight to your inbox weekly.