
Software Engineer, Infrastructure
Posted Jul 28

Posted Jul 28
This is a fully remote position, open to applicants in Turkey.
• Develop and sustain a Python-based fleet tracking system that oversees the entire lifecycle of servers, including contracting, procurement, intended use, pricing, availability, health, RMAs, and more.
• Create server management tools that automate tasks such as provisioning, health checks, GPU diagnostics, recovery, and alert notifications.
• Design and uphold metrics, dashboards, and alerts for hardware health across the fleet, addressing GPU errors, disk failures, network issues, and thermal conditions.
• Utilize AI to a significant extent in developing tools that automate alerting and recovery processes.
• Enforce and implement OS-level security measures, including hardening baselines, SELinux/AppArmor policies, SSH key management, vulnerability scanning, and compliance automation.
• Oversee and enhance distributed and local storage solutions that support model weights, checkpoints, and temporary scratch storage, including NVMe arrays, NFS, parallel file systems, and object storage.
• Optimize Linux systems for AI workloads by tuning kernel parameters, configuring NUMA topology, CPU pinning, hugepages, I/O schedulers, and optimizing the GPU driver stack (NVIDIA drivers, CUDA, container runtimes).
• Create a comprehensive suite of automated error detection and recovery mechanisms.
• Collaborate with partners to address technical challenges.
• A minimum of 3 years of experience managing large-scale bare-metal and cloud-based server fleets (100+ nodes).
• Strong proficiency in software engineering with Python; capable of writing production-grade tools rather than just scripts.
• Extensive knowledge of Linux systems, including the boot process, kernel tuning, networking, storage, systemd, cgroups, namespaces, and performance profiling.
• Significant experience with configuration management and infrastructure-as-code tools, such as Ansible, Terraform, and cloud-init.
• Solid understanding of storage technologies, including LVM, RAID, NVMe, NFS, Lustre or GPFS, as well as tuning the Linux I/O stack.
• Familiarity with hardware diagnostics and potential failure modes related to GPUs, NVMe, NICs, and memory.
• Experience in developing internal tools or dashboards for enhanced infrastructure visibility.
• Excellent communication skills and the capability to influence technical decisions within teams.
• A self-motivated individual who acts decisively, takes ownership of tasks, and consistently seeks improvement.
• Knowledge of network configuration and diagnostics (VLAN, VXLAN, ECMP, BGP, tcpdump) is a plus.
• Experience with NVIDIA GPU infrastructure, including driver management, health monitoring, DCGM, NVLink/NVSwitch diagnostics, RDMA, and InfiniBand/RoCEv2 is a plus.
• Experience with AMD GPUs is a plus.
• Familiarity with provisioning bare metal and VMs (PXE/iPXE, Kickstart, libvirt, Qemu/KVM) is a plus.
• Understanding of compliance frameworks applicable to cloud providers (SOC 2, ISO 27001) is a plus.
• Engaging and challenging work.
• Numerous opportunities for learning and professional growth.
• Regular team events and offsite gatherings.
Teleperformance
Trilon Group
Carbon60
fal
Get handpicked remote jobs straight to your inbox weekly.