Senior Site Reliability Engineer – MAAS

Posted 4 days ago

This is a fully remote position, open to applicants in Armenia.

📋 Description

• Manage and maintain extensive Linux infrastructure within Debian/Ubuntu-based bare-metal and virtualized environments.

• Take responsibility for MAAS-based bare-metal provisioning, including region/rack controllers, PXE, commissioning, cloud-init, node lifecycle, and API/CLI automation.

• Oversee and maintain production Kubernetes clusters, focusing on upgrades, node pools, networking, storage, security hardening, and troubleshooting.

• Design and sustain multi-site networking across VLANs, L2/L3 routing, bonded interfaces, VPNs, firewalls, and DNS.

• Automate infrastructure provisioning and operations through Ansible, Bash/Python, OpenTofu/Terraform, and Git-based workflows.

• Develop and sustain automated deployment workflows including PXE, Preseed, and cloud-init.

• Operate observability platforms using Prometheus, Grafana, Alertmanager, VictoriaMetrics/VictoriaLogs, or similar tools.

• Establish and enhance SLIs, SLOs, alerting, and reliability practices across infrastructure and platform services.

• Lead infrastructure incident response, troubleshooting, escalation, and post-incident improvements.

• Maintain on-call processes and operational coverage across distributed environments.

• Work closely with the hardware layer, including IPMI/Redfish, BMCs, RAID, storage, hardware diagnostics, and GPU infrastructure.

• Manage virtualization platforms such as Proxmox, KVM/libvirt, OpenStack, or VMware, including GPU passthrough when necessary.

• Build and maintain internal infrastructure tools for host discovery, configuration, IPAM, hardware health, and operational automation.

• Take ownership of infrastructure lifecycle activities including site onboarding, maintenance, decommissioning, drift detection, and operational runbooks.

• Collaborate closely with engineering and cross-functional teams to enhance reliability, resource utilization, and operational efficiency.


⛳️ Requirements

• 5+ years of practical experience in SRE, Infrastructure, Systems, or Platform Engineering.

• Advanced Linux administration skills, especially with Debian/Ubuntu.

• Significant production experience with MAAS and bare-metal provisioning.

• Expert-level, hands-on experience managing Kubernetes in production, covering cluster lifecycle, networking, storage, upgrades, and troubleshooting.

• Strong networking engineering capabilities across VLANs, L2/L3 routing, bonding, VPNs, firewalls, and DNS.

• Proficient automation skills with Ansible, Bash, and/or Python.

• Familiarity with Terraform/OpenTofu and Git-based infrastructure workflows.

• Production experience with Prometheus/Grafana or similar observability platforms.

• Background in incident response, on-call operations, monitoring, alerting, and reliability practices.

• Experience with Proxmox, KVM/libvirt, OpenStack, VMware, or related virtualization technologies.

• Knowledge of bare-metal hardware, BMCs, IPMI/Redfish, storage, and hardware troubleshooting.

• Strong understanding of distributed systems, container orchestration, and infrastructure reliability.

• Experience with infrastructure security including RBAC, firewalls, network policies, secrets management, and security hardening.

• Ability to create SOPs, runbooks, and operational processes from the ground up.

• Comfortable working independently in a fast-paced, engineering-focused environment.

• Proficient in English is required.


🏝️ Benefits

• Fully remote position with flexible working hours.

• High-impact role offering substantial technical ownership and autonomy.

• Work directly with bare-metal, Kubernetes, networking, and GPU infrastructure.

• Join an international, engineering-driven team.

• Strong emphasis on automation, reliability, and large-scale infrastructure.

• Opportunity to influence the architecture and operational foundations of a budding cloud platform.

People also viewed

Koniag Government Services1 day ago

Architect/DevOps Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
FP Markets (First Prudential Markets)1 day ago

Senior DevOps Engineer

AM flagArmenia, +4 more countriesFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Modern Campus1 day ago

Senior DevOps Engineer

US flagUnited States, +1 more countryFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
InRule1 day ago

Site Reliability Engineer

US flagUnited States, +1 more countryFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Thumbtack1 day ago

Senior Software Engineer, Site Reliability Engineering

US flagUnited States, +38 more locationsFull-timeDevOps & Site Reliability Engineer (SRE)$179.4k – $272.8k/year
ApplyView job
Thumbtack1 day ago

Senior Software Engineer, Site Reliability Engineering

CA flagCanada, +1 more countryFull-timeDevOps & Site Reliability Engineer (SRE)C$180.2k – C$233.2k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers