
Senior Site Reliability Engineer, Germany
Posted 1 day ago

Posted 1 day ago
This is a fully remote position, open to applicants in Germany.
• Ensure the reliable and secure operation of the current cplace Cloud 1.0 for customers.
• Design the architecture and operational framework of the Kubernetes-based cplace Cloud 2.0.
• Collaborate in the development of the Kubernetes platform, focusing on cluster design, networking, storage, tenant isolation, and scalability.
• Plan and manage the migration of customer environments from Cloud 1.0.
• Create reusable Terraform modules, GitOps repositories, and platform software such as self-service portals, APIs, and automation tools.
• Perform code reviews, automated testing, and implement policy as code.
• Manage and enhance Linux servers, containers, SQL databases, and Elasticsearch.
• Address bugs, findings, and capacity challenges through incident and problem management, along with post-mortem evaluations.
• Transform manual processes into versioned, tested code utilizing Ansible, Terraform, and n8n.
• Enhance monitoring, logging, and alerting practices.
• Oversee backup and recovery operations, disaster recovery, and business continuity, ensuring regular testing.
• Execute security and compliance mandates such as SOC 2, ISO 27001, and GDPR.
• Optimize cost efficiency and capacity utilization.
• Propel AI initiatives in operations, including incident and log analysis, as well as automation for runbooks and routine tasks.
• Work alongside product development teams to facilitate seamless releases.
• Represent the technical aspects of the team in customer meetings, tenders, and projects.
• Disseminate knowledge through both internal and external documentation and mentoring.
• Participate in on-call duties on a fair rotation basis.
• Bachelor's degree in computer science or a related STEM discipline, or equivalent vocational training.
• Several years of experience, ideally 5+ years, as an SRE, DevOps, or platform engineer in critical production environments.
• Strong hands-on experience with Kubernetes in production settings, covering operations, upgrades, troubleshooting, storage, and networking.
• Familiarity with at least one cloud provider; experience with AWS, GCP, Azure, and/or Hetzner Cloud is highly advantageous.
• Extensive experience with infrastructure as code using Terraform and Ansible.
• Proficiency in CI/CD, GitOps, and Git-based collaboration through pull requests and code reviews.
• Experience in writing modular, tested, and maintainable code.
• Strong skills in Linux.
• Experience with SQL databases, such as MariaDB.
• Familiarity with Elasticsearch/OpenSearch and observability stacks like Prometheus, Grafana, and Loki/ELK.
• Solid software engineering capabilities, preferably in Go or Python.
• Proficient in Bash scripting.
• Good understanding of cloud security, including network segmentation, secrets management, and WAF.
• Hands-on experience with AI tools in engineering and a solid understanding of their capabilities and limitations.
• Customer-oriented with a structured approach to work.
• Ability to articulate technical concepts clearly.
• Fluent in both German and English.
• Flexible work model with an option for remote work.
• Competitive salary package.
• 30 days of vacation annually.
• Opportunity for sabbaticals.
• Access to Wellpass.
• Availability of a job bike.
• Corporate benefits program.
• Choice of hardware equipment.
• Opportunity for creativity, co-design, and professional development.
Harris Computer
TheWhiteam
CFactory-Creations
Capgemini
Get handpicked remote jobs straight to your inbox weekly.