
Senior Site Reliability Engineer, Kubernetes
Posted Jul 14

Posted Jul 14
This is a fully remote position, open to applicants in Greece.
• Administer and uphold Linux-based infrastructure, specifically Debian and Ubuntu distributions.
• Deploy, manage, and scale Kubernetes clusters in various environments, including bare-metal, virtualized, and on-premises.
• Supervise the complete lifecycle of clusters, including upgrades, node pools, networking, storage, and security enhancements.
• Implement automation solutions for provisioning and operational tasks using Ansible, Bash/Python, and GitOps methodologies.
• Design and sustain networking architecture, covering VLANs, L2/L3 routing, VPNs, and multi-site connectivity.
• Construct automated deployment workflows, including PXE boot, Preseed, and cloud-init processes.
• Deploy and maintain observability stacks like Prometheus/Grafana, Loki, ELK, and Graylog.
• Lead incident response efforts and escalation procedures across the platform.
• Enhance system availability and minimize latency across all operational levels.
• Define and execute SLOs/SLIs across various infrastructure tiers, including physical networks, hardware, platform virtualization, and software services.
• Streamline alerting and monitoring pipelines to deliver actionable insights.
• Establish and uphold on-call schedules to guarantee coverage across different time zones.
• Develop Standard Operating Procedures (SOPs) for consistent operations and maintenance tasks.
• Coordinate physical maintenance for Policlouds, addressing periodic maintenance, hardware issues, and DC-Ops.
• Oversee virtualization and orchestration layers such as OpenStack, Proxmox, and VMware.
• Contribute to the development and maintenance of the overall architecture across all product offerings.
• Plan resources for upcoming initiatives, considering demand and growth forecasts.
• Collaborate with development teams to enhance overall quality and optimize resource usage.
• Engage with cross-functional stakeholders, including Hivenet, Policloud, and Customer Success teams.
• Extensive, hands-on experience managing Kubernetes in production settings.
• Strong network engineering expertise, including VLANs, L2/L3 routing, VPNs, and multi-site connectivity—this is crucial for the position.
• Proficient in Linux systems administration, particularly with Debian and Ubuntu.
• Solid grasp of networking principles and the ability to architect complex network designs.
• Experience in building and maintaining automation workflows using tools such as Ansible, Bash/Python, and Git-based solutions.
• Familiarity with observability stacks, including Prometheus, Grafana, ELK, Loki, or Graylog.
• Background in virtualization technologies like OpenStack, Proxmox, and VMware.
• Experience with bare-metal provisioning and Metal as a Service (MAAS).
• Strong foundation in distributed systems and container orchestration.
• Process-oriented mindset with the ability to create SOPs and operational procedures from the ground up.
• Experience in incident response, escalation protocols, and on-call rotations.
• Capability to work independently in a fast-paced, engineering-driven environment.
• Strong technical abilities coupled with alignment to team values.
• 100% remote work with flexible hours.
• High-impact role that offers autonomy and ownership.
• Collaborative and international engineering team.
• Access to a cutting-edge tech stack with a strong emphasis on reliability and automation.
The Codest
IRIUM
Sólides
Resilinc
Get handpicked remote jobs straight to your inbox weekly.