
Senior Site Reliability Engineer, Spain
Posted 1 day ago

Posted 1 day ago
This is a fully remote position, open to applicants in Spain.
• Ensure the reliable and secure operation of cplace Cloud 1.0 for clients.
• Design the architecture and operational framework for the Kubernetes-based cplace Cloud 2.0.
• Collaborate in building the Kubernetes platform, focusing on cluster design, networking, storage, tenant isolation, and scalability.
• Strategize and facilitate the migration of customer environments from Cloud 1.0.
• Create reusable Terraform modules, GitOps repositories, and platform software, including self-service portals, APIs, and automation tools.
• Conduct code reviews, automated testing, and policy as code initiatives.
• Operate and enhance Linux servers, containers, SQL databases, and Elasticsearch systems.
• Address bugs, findings, and capacity challenges through incident management, problem resolution, and post-mortem analyses.
• Transform manual processes into version-controlled, tested code utilizing Ansible, Terraform, and n8n.
• Enhance monitoring, logging, and alerting systems.
• Oversee backup and recovery, disaster recovery, and business continuity, ensuring regular testing.
• Implement security and compliance standards such as SOC 2, ISO 27001, and GDPR.
• Optimize cost efficiency and capacity utilization.
• Advance AI applications in operations, including incident and log analysis and implementing agents for runbooks and routine tasks.
• Collaborate with product development teams to ensure seamless releases.
• Represent the technical perspective of the team in customer meetings, tenders, and projects.
• Disseminate knowledge through both internal and external documentation and mentoring.
• Participate in on-call duty with a fair rotation.
• Bachelor's degree in computer science or a related STEM discipline, or equivalent vocational training.
• Several years of experience, ideally 5+ years, as an SRE, DevOps, or platform engineer in mission-critical production settings.
• Extensive hands-on experience with Kubernetes in production, covering operations, upgrades, troubleshooting, storage, and networking.
• Proficiency with at least one cloud provider; familiarity with AWS, GCP, Azure, and/or Hetzner Cloud is a significant advantage.
• In-depth experience with infrastructure as code, particularly with Terraform and Ansible.
• Familiarity with CI/CD practices, GitOps, and Git-based collaboration through pull requests and code reviews.
• Ability to write modular, tested, and maintainable code.
• Strong expertise in Linux administration.
• Experience with SQL databases such as MariaDB, Elasticsearch/OpenSearch, and observability tools like Prometheus, Grafana, and Loki/ELK.
• Solid software engineering skills, preferably in Go or Python.
• Proficient in Bash scripting.
• Good understanding of cloud security principles, including network segmentation, secrets management, and WAF.
• Practical experience with AI tools in daily engineering practices, along with an understanding of their strengths and limitations.
• Customer-centric approach, structured work style, and the ability to communicate technical concepts clearly.
• Fluent in German at least at C1 level and proficient in English.
• Flexible work model with the option for remote work.
• Opportunities for creativity, co-design, and professional growth.
• Competitive salary package.
• 30 days of vacation with the option for a sabbatical.
• Choice of hardware equipment.
Harris Computer
TheWhiteam
CFactory-Creations
Capgemini
Get handpicked remote jobs straight to your inbox weekly.