
Senior Site Reliability Engineer, SRE
Posted Jul 30

Posted Jul 30
This is a fully remote position, open to applicants in Poland.
• Collaborate with internationally distributed teams to tackle technical challenges and enhance processes.
• Design, deploy, maintain, and troubleshoot cloud and AI infrastructure solutions utilizing open source software.
• Work closely with stakeholders to gather and refine technical specifications.
• Improve system performance, reliability, and scalability.
• Diagnose, debug, and resolve intricate technical problems.
• Engage in code reviews to uphold high quality standards.
• Keep abreast of industry trends and best practices in cloud operations and development.
• Design and execute AI-driven automation throughout the DevOps lifecycle, encompassing code development and maintenance.
• Facilitate knowledge transfer to clients during the delivery phases.
• Over 5 years of professional experience in DevOps, with a strong emphasis on cloud and infrastructure technologies, including Kubernetes and/or OpenStack.
• Experience with high-performance data center processing, networking, and storage.
• Familiarity with Golang and working knowledge of additional programming languages (Python, JavaScript).
• In-depth understanding of distributed systems, microservices architecture, and CI/CD pipelines.
• Exceptional problem-solving and debugging skills across networking and storage (both hardware and software), Linux, and Kubernetes, with a focus on performance optimization and security.
• Proven capability to lead technical tasks and work collaboratively with diverse teams.
• Confident in making independent decisions when interacting directly with clients, often with minimal day-to-day oversight.
• Excellent proficiency in written and spoken English.
• Strong customer-facing communication abilities.
• Commitment to innovation, ongoing learning, and delivering high-quality outcomes.
• Willingness to travel up to 25% if necessary, including internationally.
• Preferred: Extensive experience in network and/or storage architecture.
• Experience with high-performance computing or GPU infrastructure: GPU scheduling, MIG/vGPU, RDMA/RoCE or InfiniBand fabrics, NVLink, DCGM health-checking, GPU driver/firmware lifecycle or NVIDIA AI Enterprise.
• Active participation in the open source community, including upstream contributions and conference presentations.
• Previous experience with commercial container and virtual compute infrastructure platforms such as Rancher, Openshift, and VMware.
• Join a prominent Silicon Valley leader in the cloud infrastructure sector;
• Collaborate with exceptionally passionate, talented, and engaging colleagues, assisting Fortune 500 and Global 2000 clients in implementing next-generation cloud technologies;
• Be a part of innovative, open-source advancements;
• Excel in a dynamic environment of a youthful company that values openness, collaboration, risk-taking, and continuous growth;
• Opportunities for professional development and training;
• Attend conferences and working groups;
• Enjoy company outings, happy hours, hackathons, and tech talks;
• Receive a competitive compensation package accompanied by a robust benefits plan.
CWILL
a37
GT
Sigma Software Group
Get handpicked remote jobs straight to your inbox weekly.