Principal Site Reliability Engineer

Posted 4 days ago

This is a fully remote position, open to applicants in Poland.

📋 Description

• Develop and define the compute-platform strategy and qualification criteria for x64, ARM, accelerator, and inference hardware encompassing firmware, OS, runtime, and reliability.

• Lead enhancements in provisioning and CI/CD processes, focusing on reproducible, secure, and scalable bare-metal provisioning, imaging, configuration, and infrastructure delivery pipelines.

• Promote reliability and observability by tackling systemic failures, improving telemetry and diagnostics, and establishing safe production rollout and recovery protocols.

• Examine intricate hardware/software incidents, pinpoint root causes, direct decision-making, and translate insights into engineering solutions.

• Evaluate designs, mentor engineers, set standards, and resolve cross-domain challenges among teams and vendors.

• Facilitate new hardware integration through provisioning services and offer seamless firmware upgrade strategies.

• Guarantee dependable server performance across Akamai datacenters.


⛳️ Requirements

• Proficiency in Linux systems and server platforms, including boot processes, storage, networking, hardware diagnostics, firmware, BIOS/UEFI, BMCs, and server lifecycles.

• Experience in designing and managing bare-metal provisioning, configuration management, and scalable infrastructure automation.

• Competency in Python and Bash.

• Background in creating automation, APIs, deployment pipelines, and diagnosing multi-layer failures.

• Ability to determine the source of hardware-test failures, whether from firmware, BIOS/UEFI configuration, device firmware, kernel/driver performance, or the physical testing environment.

• Expertise in observability, metrics, capacity analysis, incident root-cause evaluation, and production readiness.

• Solid understanding of x86 and ARM platforms, accelerators, and inference infrastructure, including drivers, runtimes, and compatibility.

• Practical experience with Linux virtualization technologies, such as KVM, QEMU, and libvirt.

• Familiarity with nested virtualization, CPU virtualization extensions, VM networking and storage, and Linux-level troubleshooting in virtualized settings.


🏝️ Benefits

• Health, well-being, financial, and life beyond work benefits.

• FlexBase workplace flexibility: work from home, in an office, or a combination of both.

People also viewed

FourEnergy GmbH16 hours ago

Senior DevOps Engineer – Operations

DE flagGermany OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
ICF18 hours ago

Lead DevOps Engineer

US flagVirginia OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$131.3k – $223.1k/year
ApplyView job
Mastercam23 hours ago

DevSecOps Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
C&S Informática1 day ago

DevOps Engineer – Freelance/Contract, Mid-Level/Senior

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Convene1 day ago

Support and Deployment Engineer

SA flagSaudi Arabia OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Verity Group1 day ago

SRE Engineer

BR flagBrazil OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers