Senior Manager, Validation and HPC

Posted Aug 28

This is a fully remote position, open to applicants in California, +6 more states.

πŸ“‹ Description

β€’ Oversee and guide service HPC engineering functions focused on the design, development, installation, and validation of hardware and software for customer AI HPC systems.

β€’ Spearhead the planning, execution, and performance assessment of HPC projects.

β€’ Enhance the integrity of system-services bring-up by configuring and maintaining HPC AI network and server platforms.

β€’ Lead the deployment of hardware and software by the team.

β€’ Plan, develop, and implement system-validation procedures.

β€’ Direct team activities, tests, and strategies for the implementation of customer HPC AI systems, along with custom scripts and testing protocols.

β€’ Provide support to the HPC Engineering team and collaborate with internal departments to develop and execute strategies aimed at service quality and ongoing improvement.

β€’ Mentor team members and support their career development objectives.

β€’ Establish relationships with NVIDIA leaders, customers, partners, and collaborators.

β€’ Identify, implement, and support NVIDIA AI solutions engineering while remaining informed about industry standards and innovations.

β€’ Act as the subject matter expert for customers from initial planning calls through to implementation.

β€’ Manage daily operations, provide guidance, and oversee the development of a multi-layered HPC service professional team.

β€’ Ensure the timely completion of AI HPC data-center projects.


⛳️ Requirements

β€’ A minimum of 10 years of overall experience in IT, high-performance computing, or a related field.

β€’ At least 3 years of experience in a managerial or leadership capacity.

β€’ Proven expertise in the design, configuration, and planning of HPC systems.

β€’ Strong knowledge of HPC storage solutions.

β€’ Proficient in low-latency/high-bandwidth interconnect infrastructure, including InfiniBand and Ethernet.

β€’ Expertise in HPC system software cluster management and provisioning tools such as Slurm, Salt, and xCAT.

β€’ Proficient in shared and distributed memory parallelism, including OpenMP, MPI, NCCL, and HPL.

β€’ Experience with hardware accelerators, including GPUs.

β€’ Strong scripting skills in languages like Bash, Perl, Python, or similar.

β€’ Familiarity with programming fundamentals.

β€’ Proficient in the administration, supervision, and maintenance of secure Linux/Unix operating systems, including CentOS and Solaris.

β€’ Ability to comprehend and work with large, complex systems; troubleshoot performance issues; and resolve infrastructure-related network challenges.

β€’ Expertise in managing multi-vendor hardware/software, security, and network/Internet protocols.

β€’ Bachelor's degree in computer science, information systems, or a related field, or equivalent experience.

β€’ Experience with InfiniBand.

β€’ Familiarity with GPU-focused hardware/software.

β€’ Experience with MPI.

β€’ Background in automation tooling, including Ansible, Salt, or Puppet.

β€’ Experience with Ethernet and storage technologies such as Lustre or GPFS.


🏝️ Benefits

β€’ Equity

β€’ Benefits

People also viewed

DWW Akademie11 hours ago

First-Line Manager, German & English

DE flagGermany OnlyFull-timeManager€700/month
ApplyView job
WBS19 hours ago

Career Services Manager

DE flagGermany, +1 more countryFull-timeManager
ApplyView job
ALB Conciergerie20 hours ago

Area Manager

FR flagFrance OnlyFreelanceManager€2,000 – €8,700/month
ApplyView job
ALB Conciergerie20 hours ago

City Manager, Real Estate

FR flagFrance OnlyFreelanceManager€2,000 – €8,700/month
ApplyView job
ALB Conciergerie20 hours ago

City Manager – Real Estate

FR flagFrance OnlyFreelanceManager€2,000 – €8,700/month
ApplyView job
ALB Conciergerie20 hours ago

City Manager, Real Estate

FR flagFrance OnlyFreelanceManager€2,000 – €8,700/month
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers