
Senior Site Reliability Engineer – Kubernetes
Posted Jul 15

Posted Jul 15
This is a fully remote position, open to applicants in Romania.
• Operate and manage infrastructure based on Linux (Debian/Ubuntu).
• Deploy, oversee, and scale Kubernetes clusters across bare-metal, virtualized, and on-premises environments.
• Supervise the entire cluster lifecycle, including upgrades, node pools, networking, storage, and security enhancements.
• Implement automation for provisioning and operations utilizing Ansible, Bash/Python, and GitOps workflows.
• Design and uphold networking architecture incorporating VLANs, L2/L3 routing, VPNs, and multi-site connectivity.
• Create automated deployment processes (PXE boot, Preseed, cloud-init).
• Deploy and sustain observability stacks (Prometheus/Grafana, Loki, ELK, Graylog).
• Lead incident response and escalation processes across the platform.
• Enhance system availability and minimize latency across all levels.
• Define and enforce SLOs/SLIs at various infrastructure levels (physical network/hardware, platform virtualization, software services).
• Optimize alerting and monitoring pipelines to deliver actionable insights.
• Establish and manage on-call schedules to ensure coverage across different time zones.
• Develop Standard Operating Procedures (SOPs) for consistent operations and maintenance tasks.
• Coordinate physical maintenance for Policlouds (regular maintenance, hardware issues, DC-Ops).
• Manage virtualization and orchestration layers (OpenStack, Proxmox, VMware).
• Assist in developing and maintaining the overall architecture across all products.
• Plan resources for upcoming initiatives, considering demand and growth forecasts.
• Collaborate with development teams to enhance overall quality and optimize resource usage.
• Work with cross-functional stakeholders (Hivenet, Policloud, Customer Success teams).
• Expert-level, hands-on experience operating Kubernetes in production settings.
• Strong network engineering skills (VLANs, L2/L3 routing, VPNs, multi-site connectivity) are crucial for this role.
• High proficiency in Linux systems administration (Debian/Ubuntu).
• Solid grasp of networking fundamentals and the capability to design intricate network architectures.
• Experience in building and maintaining automation workflows (Ansible, Bash/Python, Git-based).
• Familiarity with observability stacks such as Prometheus, Grafana, ELK, Loki, or Graylog.
• Background in virtualization technologies (OpenStack, Proxmox, VMware).
• Experience with bare-metal provisioning and MAAS (Metal as a Service).
• Strong understanding of distributed systems and container orchestration.
• Process-oriented mindset with the ability to create SOPs and operational procedures from the ground up.
• Experience in incident response, escalation protocols, and on-call rotations.
• Capability to work independently in a fast-paced, engineering-focused environment.
• Strong technical skills aligned with team values.
• 100% remote work with flexible hours.
• High-impact role with autonomy and ownership.
• Collaborative and international engineering team.
• Cutting-edge tech stack with a strong emphasis on reliability and automation.
The Codest
IRIUM
Sólides
Resilinc
Get handpicked remote jobs straight to your inbox weekly.