
Linux Infrastructure Engineer, Bare Metal, Storage, AI Factory Infrastructure
Posted Aug 28

Posted Aug 28
This is a fully remote position, open to applicants in Singapore.
• Design, implement, operate, and troubleshoot extensive Linux-based infrastructure.
• Construct and oversee infrastructure from the hardware level upwards, which includes servers, networking, storage, GPU clusters, and platforms optimized for AI.
• Deploy and manage bare metal servers along with Bare Metal as a Service (BMaaS) platforms.
• Create, implement, and support large-scale enterprise Linux infrastructure.
• Deploy and oversee GPU-accelerated infrastructure tailored for AI/ML workloads.
• Provide support for GPU clusters, AI training environments, and high-performance computing (HPC) workloads.
• Provision, monitor, enhance, and manage the lifecycle of GPU infrastructure.
• Manage enterprise storage and Ceph platforms, which includes capacity planning, performance tuning, and recovery from failures.
• Design and troubleshoot high-performance networking within data centers, including Layer 2 and Layer 3 infrastructure.
• Support environments requiring high availability, clustering, disaster recovery, and mission-critical production.
• Address challenges related to operating systems, hardware, GPU, networking, and storage.
• Automate operational tasks using Bash and Python.
• Develop operational documentation, runbooks, and infrastructure standards.
• Extensive senior-level experience in Linux infrastructure engineering.
• Proficient in Linux administration; experience with Ubuntu is mandatory, while Red Hat and SUSE are preferred.
• In-depth knowledge in the deployment, architecture, provisioning, and lifecycle management of bare metal servers.
• Proven experience in managing Bare Metal as a Service (BMaaS) platforms and large-scale infrastructure environments.
• Strong understanding of server hardware components, such as BIOS/UEFI, RAID controllers, firmware management, iLO/iDRAC/IPMI, NICs and SmartNICs, HBA cards, and hardware diagnostics.
• Experience in designing, implementing, and supporting large-scale enterprise Linux infrastructure.
• Familiarity with deploying and managing GPU-accelerated AI/ML infrastructure.
• Knowledge of NVIDIA GPU technologies, including A100, H100, H200, B200, or similar platforms.
• Experience with NVIDIA DGX and OEM GPU servers, including GPU provisioning, lifecycle management, monitoring, and performance optimization.
• Understanding of AI Factory architecture and its infrastructure requirements.
• Experience in supporting GPU clusters, AI training environments, and HPC workloads.
• Knowledge of GPU resource allocation and scheduling, multi-GPU systems, GPU networking, and design principles for high-bandwidth, low-latency infrastructure.
• Familiarity with CUDA, NCCL, GPUDirect Storage, NVIDIA Fabric Manager, and NVIDIA Base Command is preferred.
• Advanced skills in Linux storage administration: LVM, XFS, EXT4, NFS, iSCSI, Fibre Channel SAN, and Multipath I/O.
• Strong hands-on experience with Ceph, covering cluster architecture, MON, OSD, MDS, RBD, CephFS, RGW, along with capacity planning, performance tuning, and failure recovery.
• Experience with high-performance AI storage solutions such as WEKA, VAST Data, Dell PowerScale, Pure Storage FlashBlade, or NetApp.
• Understanding of NVMe-over-Fabrics, RDMA, GPUDirect Storage, parallel file systems, and AI data pipelines.
• Solid networking expertise, including bonding, VLANs, routing, MTU optimization, DNS, and DHCP.
• Experience with 100G/200G/400G Ethernet, RoCE, RDMA, and Spine-Leaf architectures.
• Familiarity with NVIDIA Spectrum-X, Mellanox/NVIDIA ConnectX adapters, or similar technologies.
• Strong grasp of Layer 2 and Layer 3 infrastructure design and troubleshooting.
• Experience with high availability, clustering, and disaster recovery strategies.
• Excellent troubleshooting abilities across Linux operating systems, hardware platforms, GPU infrastructure, networking, and enterprise storage.
• Proficiency in Bash and Python scripting for automation and operational efficiency.
• Experience creating operational documentation, runbooks, and infrastructure standards.
• Familiarity with Kubernetes infrastructure, virtualization platforms, Ansible, NVIDIA Base Command Manager, Slurm/HPC workload schedulers, observability platforms, DCIM tools, IPAM solutions, and exposure to AWS/Azure/hybrid cloud is considered a plus.
• Comprehensive health, dental, and vision insurance.
• Opportunities for professional development and continuous learning.
• Flexible work hours and remote work options.
• Generous vacation and paid time off policies.
• Retirement savings plans with company matching.
Conduent
CrowdStrike
Hotel Engine
Conduent
Get handpicked remote jobs straight to your inbox weekly.