
Senior Site Reliability Engineer – MAAS
Posted 4 days ago

Posted 4 days ago
This is a fully remote position, open to applicants in Armenia.
• Manage and maintain extensive Linux infrastructure within Debian/Ubuntu-based bare-metal and virtualized environments.
• Take responsibility for MAAS-based bare-metal provisioning, including region/rack controllers, PXE, commissioning, cloud-init, node lifecycle, and API/CLI automation.
• Oversee and maintain production Kubernetes clusters, focusing on upgrades, node pools, networking, storage, security hardening, and troubleshooting.
• Design and sustain multi-site networking across VLANs, L2/L3 routing, bonded interfaces, VPNs, firewalls, and DNS.
• Automate infrastructure provisioning and operations through Ansible, Bash/Python, OpenTofu/Terraform, and Git-based workflows.
• Develop and sustain automated deployment workflows including PXE, Preseed, and cloud-init.
• Operate observability platforms using Prometheus, Grafana, Alertmanager, VictoriaMetrics/VictoriaLogs, or similar tools.
• Establish and enhance SLIs, SLOs, alerting, and reliability practices across infrastructure and platform services.
• Lead infrastructure incident response, troubleshooting, escalation, and post-incident improvements.
• Maintain on-call processes and operational coverage across distributed environments.
• Work closely with the hardware layer, including IPMI/Redfish, BMCs, RAID, storage, hardware diagnostics, and GPU infrastructure.
• Manage virtualization platforms such as Proxmox, KVM/libvirt, OpenStack, or VMware, including GPU passthrough when necessary.
• Build and maintain internal infrastructure tools for host discovery, configuration, IPAM, hardware health, and operational automation.
• Take ownership of infrastructure lifecycle activities including site onboarding, maintenance, decommissioning, drift detection, and operational runbooks.
• Collaborate closely with engineering and cross-functional teams to enhance reliability, resource utilization, and operational efficiency.
• 5+ years of practical experience in SRE, Infrastructure, Systems, or Platform Engineering.
• Advanced Linux administration skills, especially with Debian/Ubuntu.
• Significant production experience with MAAS and bare-metal provisioning.
• Expert-level, hands-on experience managing Kubernetes in production, covering cluster lifecycle, networking, storage, upgrades, and troubleshooting.
• Strong networking engineering capabilities across VLANs, L2/L3 routing, bonding, VPNs, firewalls, and DNS.
• Proficient automation skills with Ansible, Bash, and/or Python.
• Familiarity with Terraform/OpenTofu and Git-based infrastructure workflows.
• Production experience with Prometheus/Grafana or similar observability platforms.
• Background in incident response, on-call operations, monitoring, alerting, and reliability practices.
• Experience with Proxmox, KVM/libvirt, OpenStack, VMware, or related virtualization technologies.
• Knowledge of bare-metal hardware, BMCs, IPMI/Redfish, storage, and hardware troubleshooting.
• Strong understanding of distributed systems, container orchestration, and infrastructure reliability.
• Experience with infrastructure security including RBAC, firewalls, network policies, secrets management, and security hardening.
• Ability to create SOPs, runbooks, and operational processes from the ground up.
• Comfortable working independently in a fast-paced, engineering-focused environment.
• Proficient in English is required.
• Fully remote position with flexible working hours.
• High-impact role offering substantial technical ownership and autonomy.
• Work directly with bare-metal, Kubernetes, networking, and GPU infrastructure.
• Join an international, engineering-driven team.
• Strong emphasis on automation, reliability, and large-scale infrastructure.
• Opportunity to influence the architecture and operational foundations of a budding cloud platform.
Koniag Government Services
FP Markets (First Prudential Markets)
Modern Campus
InRule
Get handpicked remote jobs straight to your inbox weekly.