
Senior Manager, Validation and HPC
Posted Aug 28

Posted Aug 28
This is a fully remote position, open to applicants in California, +6 more states.
β’ Oversee and guide service HPC engineering functions focused on the design, development, installation, and validation of hardware and software for customer AI HPC systems.
β’ Spearhead the planning, execution, and performance assessment of HPC projects.
β’ Enhance the integrity of system-services bring-up by configuring and maintaining HPC AI network and server platforms.
β’ Lead the deployment of hardware and software by the team.
β’ Plan, develop, and implement system-validation procedures.
β’ Direct team activities, tests, and strategies for the implementation of customer HPC AI systems, along with custom scripts and testing protocols.
β’ Provide support to the HPC Engineering team and collaborate with internal departments to develop and execute strategies aimed at service quality and ongoing improvement.
β’ Mentor team members and support their career development objectives.
β’ Establish relationships with NVIDIA leaders, customers, partners, and collaborators.
β’ Identify, implement, and support NVIDIA AI solutions engineering while remaining informed about industry standards and innovations.
β’ Act as the subject matter expert for customers from initial planning calls through to implementation.
β’ Manage daily operations, provide guidance, and oversee the development of a multi-layered HPC service professional team.
β’ Ensure the timely completion of AI HPC data-center projects.
β’ A minimum of 10 years of overall experience in IT, high-performance computing, or a related field.
β’ At least 3 years of experience in a managerial or leadership capacity.
β’ Proven expertise in the design, configuration, and planning of HPC systems.
β’ Strong knowledge of HPC storage solutions.
β’ Proficient in low-latency/high-bandwidth interconnect infrastructure, including InfiniBand and Ethernet.
β’ Expertise in HPC system software cluster management and provisioning tools such as Slurm, Salt, and xCAT.
β’ Proficient in shared and distributed memory parallelism, including OpenMP, MPI, NCCL, and HPL.
β’ Experience with hardware accelerators, including GPUs.
β’ Strong scripting skills in languages like Bash, Perl, Python, or similar.
β’ Familiarity with programming fundamentals.
β’ Proficient in the administration, supervision, and maintenance of secure Linux/Unix operating systems, including CentOS and Solaris.
β’ Ability to comprehend and work with large, complex systems; troubleshoot performance issues; and resolve infrastructure-related network challenges.
β’ Expertise in managing multi-vendor hardware/software, security, and network/Internet protocols.
β’ Bachelor's degree in computer science, information systems, or a related field, or equivalent experience.
β’ Experience with InfiniBand.
β’ Familiarity with GPU-focused hardware/software.
β’ Experience with MPI.
β’ Background in automation tooling, including Ansible, Salt, or Puppet.
β’ Experience with Ethernet and storage technologies such as Lustre or GPFS.
β’ Equity
β’ Benefits
DWW Akademie
ALB Conciergerie
ALB Conciergerie
Get handpicked remote jobs straight to your inbox weekly.