
Senior Site Reliability Engineer – Kubernetes
Posted Jul 14

Posted Jul 14
This is a fully remote position, open to applicants in Latvia.
• Manage and operate Linux-based infrastructure, specifically Debian and Ubuntu.
• Deploy, supervise, and scale Kubernetes clusters within bare-metal, virtualized, and on-premises environments.
• Oversee the entire lifecycle of clusters, including upgrades, node pools, networking, storage, and security enhancements.
• Implement automation for provisioning and operational tasks using Ansible, Bash/Python, and GitOps practices.
• Design and uphold networking architecture that encompasses VLANs, L2/L3 routing, VPNs, and multi-site connectivity.
• Create automated deployment workflows, including PXE boot, Preseed, and cloud-init.
• Deploy and sustain observability stacks like Prometheus/Grafana, Loki, ELK, and Graylog.
• Lead incident response and escalation efforts across the platform.
• Enhance system availability and decrease latency at all operational levels.
• Define and implement Service Level Objectives (SLOs) and Service Level Indicators (SLIs) across various infrastructure levels, including physical network/hardware, platform virtualization, and software services.
• Optimize alerting and monitoring pipelines to yield actionable insights.
• Establish and uphold on-call schedules to ensure coverage across different time zones.
• Develop Standard Operating Procedures (SOPs) for consistent operational and maintenance tasks.
• Coordinate physical maintenance for Policlouds, addressing periodic upkeep, hardware issues, and DC-Operations.
• Manage virtualization and orchestration layers such as OpenStack, Proxmox, and VMware.
• Assist in the development and maintenance of the overall architecture for all products.
• Plan resources for upcoming initiatives, taking into account demand and growth forecasts.
• Collaborate with development teams to enhance overall quality and optimize resource usage.
• Work in partnership with cross-functional stakeholders, including Hivenet, Policloud, and Customer Success teams.
• Extensive hands-on experience managing Kubernetes in production settings.
• Strong networking engineering capabilities, particularly in VLANs, L2/L3 routing, VPNs, and multi-site connectivity, which are crucial for the role.
• Proficient in Linux systems administration, specifically Debian and Ubuntu.
• Solid grasp of networking principles with the capability to design intricate network architectures.
• Proven experience in building and maintaining automation workflows utilizing Ansible, Bash/Python, and Git-based systems.
• Familiarity with observability stacks, including Prometheus, Grafana, ELK, Loki, or Graylog.
• Background in virtualization technologies, such as OpenStack, Proxmox, and VMware.
• Experience with bare-metal provisioning and Metal as a Service (MAAS).
• Strong understanding of distributed systems and container orchestration.
• A process-oriented approach with the ability to develop SOPs and operational procedures from the ground up.
• Experience in incident response, escalation processes, and on-call rotations.
• Capability to work independently in a fast-paced, engineering-focused environment.
• Strong technical skills aligned with team values.
• 100% remote work with flexible working hours.
• A high-impact role that offers autonomy and ownership.
• Collaborative environment within an international engineering team.
• Access to a cutting-edge tech stack with a strong emphasis on reliability and automation.
The Codest
IRIUM
Sólides
Resilinc
Get handpicked remote jobs straight to your inbox weekly.