
Lead DevOps / SRE Engineer
Posted Aug 5

Posted Aug 5
This is a fully remote position, open to applicants in Poland.
• Take full ownership of the platform's reliability, scalability, and performance from start to finish.
• Design, construct, and oversee cloud infrastructure and Kubernetes environments.
• Collaborate with engineering and business stakeholders to drive enhancements to the platform.
• Implement and sustain CI/CD pipelines along with GitOps practices.
• Empower teams through developer support, mentoring, and sharing best practices.
• Ensure robust observability through effective monitoring, logging, and tracing.
• Lead incident response efforts, conduct root cause analysis, and promote continuous improvement.
• Define and manage Service Level Indicators (SLIs), Service Level Objectives (SLOs), Service Level Agreements (SLAs), and error budgets.
• Optimize cloud expenses while upholding performance and security standards.
• Contribute to the platform's strategy, architecture, and governance.
• Foster a culture of continuous improvement and knowledge sharing.
• Develop an internal platform that is AI-ready, standardizing and simplifying software delivery.
• Enable hundreds of developers and AI-driven workflows to deliver production-ready solutions more swiftly and confidently.
• Extensive hands-on experience with cloud-native application environments, including deployment, scaling, and operational management.
• Advanced knowledge of Kubernetes, covering cluster setup, workloads, autoscaling, and ensuring high availability.
• Solid experience with AWS, focusing on resource management, security best practices, and cloud-native services.
• Practical expertise in Infrastructure as Code, preferably with Terraform.
• Experience in constructing and maintaining CI/CD pipelines.
• Strong understanding of observability practices, including monitoring, logging, and tracing.
• Proficiency in at least one programming language: Python, Go, Java, or Node.js.
• Experience in incident management, troubleshooting, and conducting post-incident analyses.
• Familiarity with SLI, SLO, and SLA frameworks, as well as error budgets.
• Experience in cost optimization and possessing a FinOps mindset.
• Willingness to participate in an on-call rotation, available 18/7, and occasionally 24/7.
• Strong communication skills in English.
• Ability to collaborate across various teams and support developer enablement.
• Proactive mindset, ownership attitude, and strong problem-solving abilities.
• Desirable experience with multi-cloud environments, architectural frameworks, test automation, FinOps practices, Data & AI ecosystems, REST/Hypermedia API design, Team Topologies, mentoring, performance testing, chaos engineering, PKI, IAM, ISO 25010, ITIL, and AI tools.
• Fully remote work opportunity within Poland.
• Occasional business travel to Germany or Zurich, typically once a quarter for 3–4 days.
• Chance to collaborate with stakeholders and engage in strategic workshops.
• Learn from challenges and improve collectively.
• Open communication and minimal hierarchical structure.
• Encouragement to share ideas regardless of one's role or level of seniority.
CWILL
a37
GT
Sigma Software Group
Get handpicked remote jobs straight to your inbox weekly.