
Systems Engineer – Core Infrastructure
Posted 1 day ago

Posted 1 day ago
This is a fully remote position, open to applicants in United Kingdom.
• Design, implement, and scale a high-performance cloud computing platform along with distributed GPU infrastructure.
• Provision, manage, and optimize large-scale bare-metal servers and specialized hardware in multi-datacenter environments.
• Architect and maintain GPU-accelerated compute clusters, including managing driver deployments, CUDA runtimes, and firmware life-cycle operations.
• Build and automate KVM and QEMU hypervisor environments alongside Kubernetes container orchestration platforms.
• Implement and sustain ultra-low-latency network fabrics, VPCs, and SDN solutions.
• Deploy and scale high-throughput distributed storage architectures tailored for data-intensive training and inference pipelines.
• Optimize the performance of storage, memory, compute, and inter-node interconnects.
• Develop infrastructure-as-code configurations using tools such as Terraform, Ansible, or custom automation scripts.
• Design monitoring, alerting, and observability frameworks for system telemetry, hardware health, and thermal efficiency.
• Engage in high-availability architecture planning, disaster recovery testing, and incident response activities.
• Assess emerging server hardware, accelerator technologies, and cloud orchestration frameworks.
• Collaborate with Software and Platform Engineering teams to create APIs and control planes on physical infrastructure.
• Maintain documentation of architecture, hardware baseline specifications, and operational runbooks.
• Work closely with Software, Architecture, and Product teams to enhance infrastructure performance, reliability, and security.
• Proven experience as a Systems Engineer.
• Extensive, hands-on expertise with Linux systems, including Debian/Ubuntu, RHEL/Rocky, and Arch variants.
• Experience in kernel tuning and low-level system performance optimization.
• Demonstrated experience managing enterprise server hardware, high-density server chassis, and GPU/accelerator platforms.
• Strong proficiency with Kubernetes, Docker, KVM, and cloud management frameworks like OpenStack or custom orchestrators.
• Expertise in automated system provisioning and configuration management utilizing Ansible, Puppet, or Chef.
• Skilled in Infrastructure-as-Code practices using Terraform.
• Comprehensive understanding of BGP, VLANs, overlay networks, firewall policies, zero-trust architectures, and multi-tenant security isolation.
• Strong analytical and troubleshooting capabilities in high-pressure live operational environments.
• Passionate about high-performance computing, open infrastructure, and scalable system design.
• Excellent communication skills with the ability to collaborate across software, architecture, and operational disciplines.
• Remote work arrangement
Veeam Software
Beacon Software
Pure Storage
Lone Wolf Technologies
Get handpicked remote jobs straight to your inbox weekly.