
Senior HPC AI Cluster Engineer
Posted 18 hours ago

Posted 18 hours ago
This is a fully remote position, open to applicants in California.
• Design, implement, and support extensive HPC/AI clusters, ensuring effective monitoring, logging, and alerting systems are in place.
• Oversee Linux job and workload scheduling alongside orchestration tools.
• Create and maintain continuous integration and delivery pipelines.
• Develop tools to automate the deployment and management of large-scale infrastructure environments.
• Automate operational monitoring and alerting, facilitating self-service resource consumption.
• Install monitoring solutions for servers, networks, and storage systems.
• Troubleshoot issues from bare metal through the operating system, software stack, and application level.
• Develop, refine, and document standard methodologies for internal teams.
• Provide support for Research & Development activities and engage in POCs/POVs aimed at future enhancements.
• Collaborate with HPC, OS, GPU compute, and systems specialists to design, develop, and launch large-scale performance platforms.
• Share insights on system design and tuning strategies for large-scale computing tasks.
• Work alongside accelerated computing and deep learning software and hardware platforms, researchers, developers, and customers to enhance workflows and create unique solutions.
• A degree in Computer Science, Engineering, or a related discipline (or equivalent experience).
• Over 8 years of relevant experience.
• Familiarity with HPC and AI solution technologies, including CPUs, GPUs, high-speed interconnects, and supporting software.
• Proficiency in job scheduling workloads and orchestration tools such as Slurm and Kubernetes.
• Strong knowledge of Windows and Linux (Red Hat/CentOS and Ubuntu) networking and internals, ACLs, OS-level security, and common protocols like TCP, DHCP, and DNS.
• Experience with storage solutions including Lustre, GPFS, and Weka.io.
• Awareness of emerging storage technologies.
• Proficient in Python programming and Bash scripting.
• Experience with automation and configuration management tools such as Jenkins, Ansible, Puppet, and Chef.
• In-depth knowledge of networking protocols like InfiniBand and Ethernet.
• Strong understanding and experience with virtual systems such as VMware, Hyper-V, KVM, or Citrix.
• Familiarity with cloud computing platforms such as AWS, Azure, and Google Cloud.
• Knowledge of CPU and/or GPU architecture.
• Familiarity with Kubernetes and container-related microservice technologies.
• Experience with GPU-centric hardware/software like DGX and CUDA.
• Experience with RDMA (InfiniBand or RoCE) fabrics.
• Competitive salaries.
• Generous benefits package.
• Equity.
CVS Health
One Impression
Volga Partners
Mercor
Get handpicked remote jobs straight to your inbox weekly.