Remotery

Linux Infrastructure Engineer – Bare Metal, Storage, AI Factory Infrastructure

Posted 2 days ago

This is a fully remote position, open to applicants in India.

📋 Description

• Design, implement, manage, and troubleshoot extensive Linux-based infrastructures.

• Construct and oversee infrastructure at the hardware level, which includes servers, networking, storage, GPU clusters, and platforms ready for AI.

• Deploy and supervise GPU-accelerated infrastructure tailored for AI/ML workloads.

• Provision, monitor, enhance, and oversee the lifecycle of GPU platforms and clusters.

• Provide support for AI training environments and high-performance computing tasks.

• Design and maintain bare metal infrastructure and BMaaS platforms.

• Manage enterprise Linux storage and Ceph platforms.

• Design and troubleshoot high-performance data center networking and Layer 2/Layer 3 infrastructure.

• Support high-availability setups, clustering, disaster recovery efforts, and mission-critical production scenarios.

• Automate operational tasks utilizing Bash and Python.

• Develop operational documentation, runbooks, and infrastructure standards.


⛳️ Requirements

• Expert-level proficiency in Linux administration; experience with Ubuntu is required, while Red Hat and SUSE are preferred.

• In-depth knowledge of bare metal server deployment, architecture, provisioning, and lifecycle management.

• Experience in operating Bare Metal as a Service (BMaaS) platforms and managing large-scale infrastructure environments.

• Strong grasp of server hardware, including BIOS/UEFI, RAID controllers, firmware management, iLO/iDRAC/IPMI, NICs and SmartNICs, HBA cards, and hardware diagnostics.

• Proven experience in designing, implementing, and supporting enterprise Linux infrastructure at scale.

• Background in deploying and managing GPU-accelerated infrastructure for AI/ML workloads.

• Familiarity with NVIDIA GPU technologies, such as A100, H100, H200, B200, or similar platforms.

• Experience with NVIDIA DGX and OEM GPU servers, including GPU provisioning, lifecycle management, monitoring, and performance optimization.

• Knowledge of AI Factory architecture and its infrastructure requirements.

• Experience in supporting GPU clusters, AI training setups, and HPC workloads.

• Understanding of GPU resource allocation and scheduling, multi-GPU systems, GPU networking, and high-bandwidth, low-latency infrastructure design.

• Familiarity with CUDA, NCCL, GPUDirect Storage, NVIDIA Fabric Manager, and NVIDIA Base Command.

• Advanced skills in Linux storage administration: LVM, XFS, EXT4, NFS, iSCSI, Fibre Channel SAN, and Multipath I/O.

• Strong hands-on experience with Ceph, covering MON, OSD, MDS, RBD, CephFS, RGW, capacity planning, performance tuning, and failure recovery.

• Experience with high-performance AI storage solutions such as WEKA, VAST Data, Dell PowerScale, Pure Storage FlashBlade, and NetApp.

• Understanding of NVMe-over-Fabrics, RDMA, GPUDirect Storage, parallel file systems, and AI data pipelines.

• Solid networking knowledge, including bonding, VLANs, routing, MTU optimization, DNS, and DHCP.

• Experience with 100G/200G/400G Ethernet, RoCE, RDMA, and spine-leaf architectures.

• Familiarity with NVIDIA Spectrum-X, Mellanox/NVIDIA ConnectX adapters, or similar technologies.

• Comprehensive understanding of Layer 2 and Layer 3 infrastructure design and troubleshooting.

• Experience in high availability, clustering, and disaster recovery.

• Strong troubleshooting abilities across Linux operating systems, hardware platforms, GPU infrastructure, networking, and enterprise storage.

• Proficient in Bash and Python scripting for automation and operational efficiency.

• Experience in creating operational documentation, runbooks, and infrastructure standards.

• Nice to have: exposure to Kubernetes infrastructure, KVM, VMware, OpenShift Virtualization, Ansible, NVIDIA Base Command Manager, Slurm, Prometheus, Grafana, OpenTelemetry, DCIM tools, IPAM solutions, and AWS, Azure, or hybrid cloud experience.


🏝️ Benefits

• Comprehensive health, dental, and vision insurance.

• Flexible work hours and remote work options.

• Opportunities for professional development and training.

• Access to cutting-edge technology and resources.

• Collaborative work environment with a focus on innovation.

People also viewed

Conduent16 hours ago

Senior Manager, Infrastructure Engineering

US flagUnited States OnlyFull-timeInfrastructure Engineer$122.4k – $159k/year
ApplyView job
Conduent1 day ago

Senior Manager, Infrastructure Engineering

US flagUnited States OnlyFull-timeInfrastructure Engineer$122.4k – $159k/year
ApplyView job
PhoenixTeam1 day ago

AWS Infrastructure Engineer – RHEL 7, RHEL 8

US flagUnited States OnlyFull-timeInfrastructure Engineer$96k – $128k/year
ApplyView job
NVIDIA1 day ago

Data Center Infrastructure Specialist

AU flagAustralia OnlyFull-timeInfrastructure Engineer
ApplyView job
Global Radiance Review1 day ago

Cloud Infrastructure Engineer

US flagUnited States OnlyFull-timeInfrastructure Engineer
ApplyView job
Mirantis1 day ago

Software Engineer, Infrastructure – Go

US flagUnited States OnlyFull-timeInfrastructure Engineer
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers