
HPC Solutions Engineer
Posted Jul 3

Posted Jul 3
This is a fully remote position, open to applicants in United States.
• Collaborate with clients during technical discovery to establish their requirements and deliverables for various use cases, assisting them in effectively utilizing distributed GPU computing resources.
• Identify and suggest the most suitable tools for each client while developing boilerplate and reference implementations that can be reused for future clients.
• Oversee the management of GPU clusters and liaise with IT to ensure the optimal performance of NVIDIA GPUs, leveraging technologies like InfiniBand for high-speed networking.
• Supervise the deployment and upkeep of machine learning environments using virtual storage solutions and distributed computing (HPC) tools such as SLURM.
• Automate the provisioning and management of infrastructure using tools like Ansible and Terraform.
• Work in partnership with data scientists and engineers to facilitate the smooth integration of ML models into production settings.
• Conduct system performance tuning and optimization to enhance throughput and minimize latency.
• Remain informed about the latest trends in machine learning technologies and HPC to apply best practices in infrastructure configuration and model development.
• Create and maintain documentation for operational procedures and system configurations.
• Over 7 years of experience in high-performance computing, distributed machine learning, GPU computing, and/or system architecture.
• Proficient in managing NVIDIA GPU environments with knowledge of GPU computing frameworks and libraries.
• Extensive experience with high-speed networking technologies, particularly InfiniBand.
• Familiarity with HPC job schedulers, preferably SLURM.
• Expertise in automating environment setup and maintenance using Ansible and Terraform.
• Proven ability in deploying and managing virtual storage solutions.
• Strong programming skills in Python and familiarity with machine learning libraries and frameworks.
• Exceptional problem-solving, communication, and teamwork abilities.
• Engage with a wide range of hardware configurations and locations available in the industry, alongside cutting-edge GPU use cases.
• This position is fully remote, fostering a culture of high accountability and autonomy.
• Competitive salary, equity, and benefits package offered.
• Flexible PTO policy in place.
3Core Systems, Inc
Fortress Information Security
BizFirst LLC
Momentus Technologies
Get handpicked remote jobs straight to your inbox weekly.