
Support Engineer, GPU Infrastructure
Posted Sep 15

Posted Sep 15
This is a fully remote position, open to applicants in Florida.
• Diagnose Linux server issues comprehensively, encompassing boot and network boot failures, kernel and driver malfunctions, filesystems, storage constraints, services, memory and CPU behavior, and system instability.
• Identify hardware failures using out-of-band management, sensor data, POST and boot errors, SMART data, and vendor diagnostics.
• Isolate server-side network issues related to NICs, drivers, VLANs, addressing, routing, MTU, DNS, DHCP, bonding, link state, and packet captures.
• Troubleshoot NVIDIA GPU servers, focusing on GPU availability, thermal throttling, driver and VBIOS mismatches, PCIe, XID errors, and host-level conditions.
• Distinguish between hardware, OS, network, application, and configuration issues prior to escalation.
• Manage incidents until resolution or seamless handoff, determining severity based on blast radius and identifying related tickets.
• Escalate issues to engineering with comprehensive evidence packs and engage in root cause analysis and post-incident reviews.
• Collaborate with data center partners for remote hands, reboots, cabling, optics checks, component replacement, and physical inspections.
• Initiate and monitor hardware RMAs with OEMs throughout the replacement process and validate repaired or replaced equipment.
• Maintain precise asset records while supporting server turn-ups, migrations, and decommissions.
• Communicate clear causes and timeline updates to technically adept customers.
• Draft and enhance runbooks, operational procedures, categories, and closure reasons.
• Develop Python, Bash, or similar scripts and tools for health checks, data collection, and routine operations.
• Contribute to infrastructure-as-code, configuration management, monitoring, alerting, and enhancements to support workflows.
• Adhere to a defined shift schedule, participate in an escalation rotation, and provide written handoffs at the end of each shift.
• A minimum of three years of experience in supporting production servers, data center infrastructure, or bare metal and cloud environments.
• Proficient hands-on Linux troubleshooting skills, including logs, dmesg, systemd, storage tools, and network utilities.
• Familiarity with server hardware, covering CPU and memory, storage and filesystems, RAID, PCIe, NICs, power, BIOS and UEFI, firmware, and drivers.
• Experience with out-of-band management tools such as IPMI, Redfish, iDRAC, iLO, or similar.
• Working knowledge of TCP/IP and the ability to ascertain if a problem lies within the host or network.
• Strong fault-domain reasoning and sound judgment in production environments.
• Experience with ticketing, monitoring, incident management, or infrastructure management systems.
• Proficient written communication skills in English.
• Desirable but not mandatory: Experience with NVIDIA GPU servers at scale; familiarity with CUDA, NCCL, NVLink, or DCGM; exposure to HPC or AI training environments; knowledge of InfiniBand or high-performance Ethernet; experience with Dell, HPE, Supermicro, or Lenovo platforms; understanding of NVMe, ZFS, Ceph, or distributed storage; experience with Prometheus, Grafana, or similar observability tools; knowledge of NetBox or other infrastructure and asset registers; familiarity with Ansible, Terraform, or configuration management; Git-based infrastructure workflows; and experience with optics, transceivers, DAC, or AOC cabling.
• Defined shift schedule established prior to commencement.
• Participation in an escalation rotation for high-severity issues occurring outside of shift hours.
• Written handoff provided at the conclusion of each shift.
• Opportunity to influence the design of a newly established support function.
J-Mack Technologies, LLC
Truelogic Software
JumpCloud
Vempra
Get handpicked remote jobs straight to your inbox weekly.