
Linux Infrastructure Engineer – Bare Metal, Storage, AI Factory Infrastructure
Posted Aug 28

Posted Aug 28
This is a fully remote position, open to applicants in Romania.
• Design, implement, manage, and resolve issues related to extensive Linux-based infrastructure tailored for enterprise workloads and AI/ML environments.
• Construct and oversee infrastructure from the ground up, encompassing servers, networking, storage, GPU clusters, and platforms optimized for AI.
• Deploy, provision, manage, and oversee the lifecycle of bare metal servers and Bare Metal as a Service (BMaaS) platforms.
• Implement and manage GPU-accelerated infrastructure, GPU clusters, and environments for AI training.
• Provide support for high-performance computing (HPC) workloads and AI Factory environments.
• Manage enterprise-level Linux storage, Ceph clusters, and high-performance storage solutions for AI.
• Design and troubleshoot data center networking with high bandwidth and low latency.
• Ensure the stability of mission-critical production environments through high availability, clustering, and disaster recovery strategies.
• Diagnose and resolve issues related to operating systems, hardware, GPU, networking, and storage.
• Automate operational processes using Bash and Python scripting.
• Develop operational documentation, runbooks, and standards for infrastructure.
• Collaborate with a dedicated DevOps team, with a focus on infrastructure rather than CI/CD or application delivery.
• Expert-level proficiency in Linux administration; experience with Ubuntu is essential, while knowledge of Red Hat and SUSE is preferred.
• In-depth experience in deploying, architecting, provisioning, and managing the lifecycle of bare metal servers.
• Proven experience in operating Bare Metal as a Service (BMaaS) platforms and managing large-scale infrastructure environments.
• Comprehensive understanding of server hardware, including BIOS/UEFI, RAID controllers, firmware management, iLO/iDRAC/IPMI, NICs and SmartNICs, HBA cards, and hardware diagnostics.
• Experience in designing, implementing, and supporting enterprise Linux infrastructure at scale.
• Familiarity with deploying and managing GPU-accelerated infrastructure for AI/ML workloads.
• Knowledge of NVIDIA GPU technologies, such as A100, H100, H200, B200, or similar platforms.
• Experience with NVIDIA DGX and OEM GPU servers, including GPU provisioning, lifecycle management, monitoring, and performance optimization.
• Understanding of AI Factory architecture, GPU clusters, AI training environments, and HPC workloads.
• Knowledge of GPU resource allocation and scheduling, multi-GPU systems, GPU networking, and the design of high-bandwidth, low-latency infrastructure.
• Familiarity with CUDA, NCCL, GPUDirect Storage, NVIDIA Fabric Manager, and NVIDIA Base Command.
• Advanced skills in Linux storage administration, including LVM, XFS, EXT4, NFS, iSCSI, Fibre Channel SAN, and Multipath I/O.
• Significant hands-on experience with Ceph, including cluster architecture, MON, OSD, MDS, RBD, CephFS, RGW, capacity planning, performance tuning, and failure recovery.
• Experience with high-performance AI storage solutions such as WEKA, VAST Data, Dell PowerScale, Pure Storage FlashBlade, and NetApp.
• Understanding of NVMe-over-Fabrics (NVMe-oF), RDMA, GPUDirect Storage, parallel file systems, and AI data pipelines.
• Strong networking expertise, including bonding, VLANs, routing, MTU optimization, DNS, and DHCP.
• Experience with 100G/200G/400G Ethernet, RoCE, RDMA, and Spine-Leaf architectures.
• Familiarity with NVIDIA Spectrum-X, Mellanox/NVIDIA ConnectX adapters, or equivalent technologies.
• Solid understanding of Layer 2 and Layer 3 infrastructure design and troubleshooting.
• Experience with high availability, clustering, and disaster recovery practices.
• Excellent troubleshooting abilities across Linux operating systems, hardware platforms, GPU infrastructure, networking, and enterprise storage.
• Proficiency in Bash and Python scripting for automation and operational efficiency.
• Experience in creating operational documentation, runbooks, and infrastructure standards.
• Familiarity with Kubernetes infrastructure, virtualization platforms, Ansible, NVIDIA Base Command Manager, Slurm/HPC schedulers, observability platforms, DCIM tools, IPAM solutions, and exposure to AWS/Azure/hybrid cloud environments is a plus.
• Competitive salary and performance-based bonuses.
• Comprehensive health, dental, and vision insurance plans.
• Opportunities for professional development and continuous learning.
• Flexible working hours and remote work options.
• A collaborative and innovative work environment.
Conduent
CrowdStrike
Hotel Engine
Conduent
Get handpicked remote jobs straight to your inbox weekly.