Linux Infrastructure Engineer, Bare Metal, Storage, AI Factory Infrastructure

Posted Aug 28

This is a fully remote position, open to applicants in Singapore.

📋 Description

• Design, implement, operate, and troubleshoot extensive Linux-based infrastructure.

• Construct and oversee infrastructure from the hardware level upwards, which includes servers, networking, storage, GPU clusters, and platforms optimized for AI.

• Deploy and manage bare metal servers along with Bare Metal as a Service (BMaaS) platforms.

• Create, implement, and support large-scale enterprise Linux infrastructure.

• Deploy and oversee GPU-accelerated infrastructure tailored for AI/ML workloads.

• Provide support for GPU clusters, AI training environments, and high-performance computing (HPC) workloads.

• Provision, monitor, enhance, and manage the lifecycle of GPU infrastructure.

• Manage enterprise storage and Ceph platforms, which includes capacity planning, performance tuning, and recovery from failures.

• Design and troubleshoot high-performance networking within data centers, including Layer 2 and Layer 3 infrastructure.

• Support environments requiring high availability, clustering, disaster recovery, and mission-critical production.

• Address challenges related to operating systems, hardware, GPU, networking, and storage.

• Automate operational tasks using Bash and Python.

• Develop operational documentation, runbooks, and infrastructure standards.


⛳️ Requirements

• Extensive senior-level experience in Linux infrastructure engineering.

• Proficient in Linux administration; experience with Ubuntu is mandatory, while Red Hat and SUSE are preferred.

• In-depth knowledge in the deployment, architecture, provisioning, and lifecycle management of bare metal servers.

• Proven experience in managing Bare Metal as a Service (BMaaS) platforms and large-scale infrastructure environments.

• Strong understanding of server hardware components, such as BIOS/UEFI, RAID controllers, firmware management, iLO/iDRAC/IPMI, NICs and SmartNICs, HBA cards, and hardware diagnostics.

• Experience in designing, implementing, and supporting large-scale enterprise Linux infrastructure.

• Familiarity with deploying and managing GPU-accelerated AI/ML infrastructure.

• Knowledge of NVIDIA GPU technologies, including A100, H100, H200, B200, or similar platforms.

• Experience with NVIDIA DGX and OEM GPU servers, including GPU provisioning, lifecycle management, monitoring, and performance optimization.

• Understanding of AI Factory architecture and its infrastructure requirements.

• Experience in supporting GPU clusters, AI training environments, and HPC workloads.

• Knowledge of GPU resource allocation and scheduling, multi-GPU systems, GPU networking, and design principles for high-bandwidth, low-latency infrastructure.

• Familiarity with CUDA, NCCL, GPUDirect Storage, NVIDIA Fabric Manager, and NVIDIA Base Command is preferred.

• Advanced skills in Linux storage administration: LVM, XFS, EXT4, NFS, iSCSI, Fibre Channel SAN, and Multipath I/O.

• Strong hands-on experience with Ceph, covering cluster architecture, MON, OSD, MDS, RBD, CephFS, RGW, along with capacity planning, performance tuning, and failure recovery.

• Experience with high-performance AI storage solutions such as WEKA, VAST Data, Dell PowerScale, Pure Storage FlashBlade, or NetApp.

• Understanding of NVMe-over-Fabrics, RDMA, GPUDirect Storage, parallel file systems, and AI data pipelines.

• Solid networking expertise, including bonding, VLANs, routing, MTU optimization, DNS, and DHCP.

• Experience with 100G/200G/400G Ethernet, RoCE, RDMA, and Spine-Leaf architectures.

• Familiarity with NVIDIA Spectrum-X, Mellanox/NVIDIA ConnectX adapters, or similar technologies.

• Strong grasp of Layer 2 and Layer 3 infrastructure design and troubleshooting.

• Experience with high availability, clustering, and disaster recovery strategies.

• Excellent troubleshooting abilities across Linux operating systems, hardware platforms, GPU infrastructure, networking, and enterprise storage.

• Proficiency in Bash and Python scripting for automation and operational efficiency.

• Experience creating operational documentation, runbooks, and infrastructure standards.

• Familiarity with Kubernetes infrastructure, virtualization platforms, Ansible, NVIDIA Base Command Manager, Slurm/HPC workload schedulers, observability platforms, DCIM tools, IPAM solutions, and exposure to AWS/Azure/hybrid cloud is considered a plus.


🏝️ Benefits

• Comprehensive health, dental, and vision insurance.

• Opportunities for professional development and continuous learning.

• Flexible work hours and remote work options.

• Generous vacation and paid time off policies.

• Retirement savings plans with company matching.

People also viewed

Conduent9 hours ago

Windows Infrastructure Engineer

US flagUnited States OnlyFull-timeInfrastructure Engineer$85.5k – $111k/year
ApplyView job
CrowdStrike10 hours ago

Senior Infrastructure Engineer, TechOps CICD

US flagCalifornia OnlyFull-timeInfrastructure Engineer$140k – $215k/year
ApplyView job
Hotel Engine13 hours ago

Senior Software Engineer, Infrastructure

US flagUnited States OnlyFull-timeInfrastructure Engineer$135.2k – $187k/year
ApplyView job
Conduent1 day ago

Windows Infrastructure Engineer

US flagMaryland OnlyFull-timeInfrastructure Engineer$85.5k – $111k/year
ApplyView job
HavocAI1 day ago

Data and ML Infrastructure Engineer

US flagRhode Island OnlyFull-timeInfrastructure Engineer$150k – $185k/year
ApplyView job
Software Mind1 day ago

Senior Python Developer, Infrastructure

RO flagRomania, +1 more countryFull-timeInfrastructure Engineer
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers