
Linux Infrastructure Engineer – Bare Metal, Storage, AI Factory Infrastructure
Posted Aug 28

Posted Aug 28
This is a fully remote position, open to applicants in Poland.
• Design, implement, operate, and troubleshoot extensive Linux-based infrastructure.
• Support both traditional enterprise workloads and contemporary AI/ML environments.
• Construct and oversee infrastructure from the hardware level upwards, which includes servers, networking, storage, GPU clusters, and AI-ready platforms.
• Deploy and manage GPU-accelerated infrastructure tailored for AI/ML workloads.
• Provide support for GPU clusters, AI training environments, and high-performance computing (HPC) tasks.
• Manage Bare Metal as a Service (BMaaS) platforms and expansive infrastructure settings.
• Design, implement, and maintain enterprise Linux infrastructure on a large scale.
• Administer and enhance enterprise storage systems and high-performance AI storage platforms.
• Oversee Ceph cluster architecture, including capacity planning, performance tuning, and failure recovery.
• Design and troubleshoot high-performance data center networking, along with Layer 2/Layer 3 infrastructure.
• Ensure support for high availability, clustering, disaster recovery, and mission-critical production environments.
• Conduct troubleshooting across Linux operating systems, hardware, GPU infrastructure, networking, and storage components.
• Utilize Bash and Python for automation and improving operational efficiency.
• Develop operational documentation, runbooks, and infrastructure standards.
• Expert-level Linux administration is required; proficiency in Ubuntu is essential, with Red Hat and SUSE being preferred.
• Extensive knowledge in bare metal server deployment, architecture, provisioning, and lifecycle management.
• Experience in operating Bare Metal as a Service (BMaaS) platforms and large-scale infrastructure environments.
• Strong comprehension of server hardware, encompassing BIOS/UEFI, RAID controllers, firmware management, iLO/iDRAC/IPMI, NICs and SmartNICs, HBA cards, and hardware diagnostics.
• Proven experience in designing, implementing, and supporting enterprise Linux infrastructure at scale.
• Familiarity with deploying and managing GPU-accelerated infrastructure for AI/ML workloads.
• Understanding of NVIDIA GPU technologies, such as A100, H100, H200, B200, or similar platforms.
• Proficiency with NVIDIA DGX and OEM GPU servers, including GPU provisioning, lifecycle management, monitoring, and performance optimization.
• Knowledge of AI Factory architecture and its infrastructure requirements.
• Experience providing support for GPU clusters, AI training environments, and HPC workloads.
• Familiarity with GPU resource allocation and scheduling, multi-GPU systems, GPU networking, and high-bandwidth, low-latency infrastructure design.
• Understanding of CUDA, NCCL, GPUDirect Storage, NVIDIA Fabric Manager, and NVIDIA Base Command.
• Advanced skills in Linux storage administration, including LVM, XFS, EXT4, NFS, iSCSI, Fibre Channel, SAN, and Multipath I/O.
• Significant hands-on experience with Ceph, covering MON, OSD, MDS, RBD, CephFS, RGW, capacity planning, performance tuning, and failure recovery.
• Familiarity with high-performance AI storage solutions like WEKA, VAST Data, Dell PowerScale, Pure Storage FlashBlade, and NetApp.
• Understanding of NVMe-over-Fabrics (NVMe-oF), RDMA, GPUDirect Storage, parallel file systems, and AI data pipelines.
• Strong networking expertise, including bonding, VLANs, routing, MTU optimization, DNS, and DHCP.
• Experience with 100G/200G/400G Ethernet, RoCE, RDMA, and Spine-Leaf architectures.
• Familiarity with NVIDIA Spectrum-X, Mellanox/NVIDIA ConnectX adapters, or similar technologies.
• Solid understanding of Layer 2 and Layer 3 infrastructure design and troubleshooting.
• Experience with high availability, clustering, and disaster recovery mechanisms.
• Strong troubleshooting capabilities across Linux operating systems, hardware platforms, GPU infrastructure, networking, and enterprise storage.
• Proficient in Bash and Python scripting for automation and operational efficiency.
• Experience in creating operational documentation, runbooks, and infrastructure standards.
• Familiarity with Kubernetes infrastructure, virtualization platforms, Ansible, Slurm/HPC schedulers, observability platforms, DCIM tools, IPAM solutions, and AWS/Azure or hybrid cloud is advantageous.
• Note: This role does not focus on DevOps; experience with CI/CD, Terraform, GitOps, application delivery, cloud-only administration, or software development is not required.
• Comprehensive health and wellness benefits.
• Opportunities for professional growth and development.
• Flexible working arrangements to support work-life balance.
• Engaging work environment with a focus on innovation and teamwork.
Conduent
CrowdStrike
Hotel Engine
Conduent
Get handpicked remote jobs straight to your inbox weekly.