
Principal Site Reliability Engineer
Posted 4 days ago

Posted 4 days ago
This is a fully remote position, open to applicants in Poland.
• Develop and define the compute-platform strategy and qualification criteria for x64, ARM, accelerator, and inference hardware encompassing firmware, OS, runtime, and reliability.
• Lead enhancements in provisioning and CI/CD processes, focusing on reproducible, secure, and scalable bare-metal provisioning, imaging, configuration, and infrastructure delivery pipelines.
• Promote reliability and observability by tackling systemic failures, improving telemetry and diagnostics, and establishing safe production rollout and recovery protocols.
• Examine intricate hardware/software incidents, pinpoint root causes, direct decision-making, and translate insights into engineering solutions.
• Evaluate designs, mentor engineers, set standards, and resolve cross-domain challenges among teams and vendors.
• Facilitate new hardware integration through provisioning services and offer seamless firmware upgrade strategies.
• Guarantee dependable server performance across Akamai datacenters.
• Proficiency in Linux systems and server platforms, including boot processes, storage, networking, hardware diagnostics, firmware, BIOS/UEFI, BMCs, and server lifecycles.
• Experience in designing and managing bare-metal provisioning, configuration management, and scalable infrastructure automation.
• Competency in Python and Bash.
• Background in creating automation, APIs, deployment pipelines, and diagnosing multi-layer failures.
• Ability to determine the source of hardware-test failures, whether from firmware, BIOS/UEFI configuration, device firmware, kernel/driver performance, or the physical testing environment.
• Expertise in observability, metrics, capacity analysis, incident root-cause evaluation, and production readiness.
• Solid understanding of x86 and ARM platforms, accelerators, and inference infrastructure, including drivers, runtimes, and compatibility.
• Practical experience with Linux virtualization technologies, such as KVM, QEMU, and libvirt.
• Familiarity with nested virtualization, CPU virtualization extensions, VM networking and storage, and Linux-level troubleshooting in virtualized settings.
• Health, well-being, financial, and life beyond work benefits.
• FlexBase workplace flexibility: work from home, in an office, or a combination of both.
FourEnergy GmbH
ICF
Mastercam
C&S Informática
Get handpicked remote jobs straight to your inbox weekly.