Linux Infrastructure Engineer – Bare Metal, Storage, AI Factory Infrastructure

Posted Aug 28

This is a fully remote position, open to applicants in Serbia.

📋 Description

• Design, implement, manage, and resolve issues within extensive Linux-based infrastructure.

• Support both traditional enterprise workloads and contemporary AI/ML environments.

• Construct and oversee infrastructure from the hardware level up, encompassing servers, networking, storage, GPU clusters, and AI-capable platforms.

• Deploy and maintain GPU-accelerated infrastructure tailored for AI/ML workloads.

• Provide support for GPU clusters, AI training environments, and high-performance computing (HPC) workloads.

• Manage and operate bare metal infrastructure along with Bare Metal as a Service (BMaaS) platforms.

• Oversee and maintain enterprise storage solutions, including Ceph and high-performance AI storage platforms.

• Design and troubleshoot high-performance data center networking along with Layer 2 and Layer 3 infrastructure.

• Ensure support for high-availability, clustering, disaster recovery, and critical production environments.

• Diagnose issues related to Linux, hardware, GPU, networking, and storage.

• Automate operational processes utilizing Bash and Python.

• Develop operational documentation, runbooks, and infrastructure standards.

• Work autonomously within a dedicated DevOps team; this position is not primarily focused on DevOps.


⛳️ Requirements

• Advanced expertise in Linux administration, with Ubuntu as a requirement; Red Hat and SUSE as preferred.

• Extensive knowledge in bare metal server deployment, architecture, provisioning, and lifecycle management.

• Experience in operating Bare Metal as a Service (BMaaS) platforms and managing vast infrastructure environments.

• Strong comprehension of server hardware, including BIOS/UEFI, RAID controllers, firmware management, iLO/iDRAC/IPMI, NICs and SmartNICs, HBA cards, and hardware diagnostics.

• Proven experience in designing, implementing, and maintaining enterprise Linux infrastructure at scale.

• Proficiency in deploying and managing GPU-accelerated infrastructure for AI/ML workloads.

• Knowledge of NVIDIA GPU technologies, such as A100, H100, H200, B200, or similar platforms.

• Familiarity with NVIDIA DGX and OEM GPU servers, GPU provisioning, lifecycle management, monitoring, and performance enhancement.

• Understanding of AI Factory architecture and infrastructure needs.

• Experience in supporting GPU clusters, AI training environments, and HPC workloads.

• Knowledge of GPU resource allocation and scheduling, multi-GPU systems, GPU networking, and high-bandwidth/low-latency infrastructure.

• Acquainted with CUDA, NCCL, GPUDirect Storage, NVIDIA Fabric Manager, and NVIDIA Base Command.

• Advanced skills in Linux storage administration: LVM, XFS, EXT4, NFS, iSCSI, Fibre Channel SAN, and Multipath I/O.

• Solid hands-on experience with Ceph, including MON, OSD, MDS, RBD, CephFS, RGW, capacity planning, performance tuning, and failure recovery.

• Experience with high-performance AI storage solutions such as WEKA, VAST Data, Dell PowerScale, Pure Storage FlashBlade, and NetApp.

• Understanding of NVMe-over-Fabrics, RDMA, GPUDirect Storage, parallel file systems, and AI data pipelines.

• Strong networking expertise, including bonding, VLANs, routing, MTU optimization, DNS, and DHCP.

• Familiarity with 100G/200G/400G Ethernet, RoCE, RDMA, and Spine-Leaf architectures.

• Knowledge of NVIDIA Spectrum-X, Mellanox/NVIDIA ConnectX adapters, or similar technologies.

• Strong understanding of Layer 2 and Layer 3 infrastructure design and troubleshooting.

• Experience with high availability, clustering, and disaster recovery mechanisms.

• Excellent troubleshooting skills across Linux operating systems, hardware platforms, GPU infrastructure, networking, and enterprise storage.

• Proficiency in Bash and Python scripting to enhance automation and operational efficiency.

• Experience in crafting operational documentation, runbooks, and infrastructure standards.

• Familiarity with Kubernetes infrastructure, virtualization platforms, Ansible, NVIDIA Base Command Manager, Slurm/HPC schedulers, observability platforms, DCIM tools, IPAM solutions, and AWS/Azure or hybrid cloud environments is a plus.


🏝️ Benefits

• Competitive salary and performance-based bonuses.

• Comprehensive health, dental, and vision insurance.

• Flexible working hours and remote work options.

• Opportunities for professional development and training.

• Collaborative and dynamic work environment.

People also viewed

VALR1 day ago

Senior Infrastructure Engineer

ZA flagSouth Africa OnlyFull-timeInfrastructure Engineer
ApplyView job
First Circle1 day ago

Senior Infrastructure Engineer

HK flagHong Kong, +4 more countriesFull-timeInfrastructure Engineer
ApplyView job
crewAI1 day ago

Software Engineer, Infrastructure, Reliability

US flagUnited States OnlyFull-timeInfrastructure Engineer
ApplyView job
Sezzle1 day ago

Principal Infrastructure Engineer

AR flagArgentina OnlyFull-timeInfrastructure Engineer$12.5k – $20.8k/month
ApplyView job
Sezzle1 day ago

Principal Infrastructure Engineer

MX flagMexico OnlyFull-timeInfrastructure Engineer$12.5k – $20.8k/month
ApplyView job
Sezzle1 day ago

Principal Infrastructure Engineer

BR flagBrazil OnlyFull-timeInfrastructure Engineer$12.5k – $20.8k/month
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers