
Senior HPC AI Cluster Engineer
Posted Sep 17

Posted Sep 17
This is a fully remote position, open to applicants in Switzerland, +4 more countries.
• Design, implement, and sustain extensive HPC/AI clusters featuring monitoring, logging, and alerting capabilities.
• Oversee Linux job/workload scheduling and orchestration tools.
• Create and uphold continuous integration and delivery pipelines.
• Develop automation tools for the deployment and management of large-scale infrastructure environments.
• Automate operational monitoring and alerting to facilitate self-service resource consumption.
• Implement monitoring solutions for servers, networks, and storage systems.
• Perform troubleshooting from the bare metal level through the operating system, software stack, and application layers.
• Develop, refine, and document best practices for internal teams.
• Assist with research and development initiatives and engage in proofs of concept and value.
• Offer insights into system design and tuning for large-scale computational tasks.
• Collaborate with scientific researchers, developers, customers, and specialists in HPC, OS, GPU computing, and systems to architect, develop, and launch large-scale performance platforms.
• A degree in Computer Science, Engineering, or a related discipline.
• Over 8 years of professional experience.
• Proficient understanding of HPC and AI solution technologies, including CPUs, GPUs, high-speed interconnects, and supporting software.
• Experience with job scheduling workloads and orchestration tools like Slurm and Kubernetes.
• Strong knowledge of Windows and Linux operating systems, particularly Red Hat/CentOS and Ubuntu.
• Familiarity with networking concepts, sockets, firewalld, iptables, Wireshark, ACLs, OS-level security measures, TCP, DHCP, and DNS.
• Experience with storage solutions including Lustre, GPFS, and Weka.io.
• Proficient in Python programming and Bash scripting.
• Familiarity with automation tools such as Jenkins, Ansible, Puppet, or Chef.
• In-depth knowledge of InfiniBand and Ethernet networking protocols.
• Experience with virtual systems like VMware, Hyper-V, KVM, or Citrix.
• Familiarity with cloud platforms such as AWS, Azure, or Google Cloud.
• Preferred: knowledge of CPU and/or GPU architecture.
• Preferred: understanding of Kubernetes and container-related microservice technologies.
• Preferred: experience with GPU-oriented hardware/software like DGX and CUDA.
• Preferred: experience with RDMA fabrics, including InfiniBand or RoCE.
• Equal opportunity employer.
• Reasonable accommodations for individuals with disabilities to engage in the application or interview process, perform essential job functions, and receive other employment-related benefits and privileges.
Mercor
Mercor
Mercor
Mercor
Get handpicked remote jobs straight to your inbox weekly.