
Infrastructure Solutions Architect – OEM
Posted 9 hours ago

Posted 9 hours ago
This is a fully remote position, open to applicants in Texas, +1 more state.
• Act as the primary technical liaison for OEM Federal implementations of NVIDIA GPU-accelerated platforms, such as PCIe GPUs, HGX, and MGX systems.
• Resolve critical issues pertaining to OEM Federal deployments.
• Manage customer concerns from start to finish, focusing on platform, firmware, and workload challenges.
• Facilitate problem resolution by collaborating with OEM support tiers and NVIDIA engineering, providing logs and reproduction details to ascertain root causes.
• Maintain communication with customers until issues are resolved.
• Develop and finalize diagnostic, RMA eligibility, and replacement protocols for Federal deployments.
• Manage scenarios where malfunctioning hardware and logs are required to stay on-site at customer locations.
• Navigate FedRAMP, IL5/IL6, and air-gapped SCIF environments effectively.
• Assist clients in adhering to sanitization and chain-of-custody standards for hardware returns.
• Collaborate with OEM Federal support and services teams to minimize blocking issues and avoid unnecessary dispatches.
• Document failure signatures to enhance the Federal knowledge base.
• Utilize AI tools where applicable within the environment.
• Must be a U.S. citizen.
• Must have and maintain an active U.S. Government TS/SCI clearance.
• Some Intelligence Community customer assignments may necessitate a current U.S. Government-administered Counterintelligence Scope or Full Scope Polygraph.
• BS or MS in Computer Engineering, Electrical Engineering, Computer Science, or equivalent experience.
• Over 5 years of engineering experience with multi-GPU platforms.
• At least 3 years in a customer-facing support, field engineering, or blocking issue role.
• Strong expertise in system software, covering firmware, BIOS, kernel, drivers, and operating systems.
• Capability to troubleshoot, enhance, and tailor Linux environments for AI/ML workloads.
• In-depth knowledge of data center infrastructure, including x86/ARM systems, high-performance storage, and low-latency networking (InfiniBand/RDMA).
• Excellent communication and organizational skills at a professional level.
• Ability to adapt to the technical proficiency of the audience, remain composed in challenging situations, and drive issues to resolution.
• Willingness to travel up to 35% to customer, partner, and data center sites, including secure facilities.
• Availability for occasional weekend and holiday coverage.
• Proven experience with Dell PowerEdge GPU servers and Dell management software (OpenManage, APEX).
• Familiarity with Dell support and blocking-issue procedures.
• Direct experience in supporting DoD, Intelligence Community, or Civilian agency deployments.
• Proficient in Python.
• Proficient in C/C++ for platform OS, firmware, and driver development.
• Experience with Docker, Kubernetes, Slurm, NCCL, and MPI.
• Experience assessing the performance of distributed GPU-accelerated workloads.
• Equity.
• Benefits.
OpenText
Connection
Calix
Grafana Labs
Get handpicked remote jobs straight to your inbox weekly.