Senior HPC AI Cluster Engineer

Posted Sep 17

This is a fully remote position, open to applicants in Switzerland, +4 more countries.

📋 Description

• Design, implement, and sustain extensive HPC/AI clusters featuring monitoring, logging, and alerting capabilities.

• Oversee Linux job/workload scheduling and orchestration tools.

• Create and uphold continuous integration and delivery pipelines.

• Develop automation tools for the deployment and management of large-scale infrastructure environments.

• Automate operational monitoring and alerting to facilitate self-service resource consumption.

• Implement monitoring solutions for servers, networks, and storage systems.

• Perform troubleshooting from the bare metal level through the operating system, software stack, and application layers.

• Develop, refine, and document best practices for internal teams.

• Assist with research and development initiatives and engage in proofs of concept and value.

• Offer insights into system design and tuning for large-scale computational tasks.

• Collaborate with scientific researchers, developers, customers, and specialists in HPC, OS, GPU computing, and systems to architect, develop, and launch large-scale performance platforms.


⛳️ Requirements

• A degree in Computer Science, Engineering, or a related discipline.

• Over 8 years of professional experience.

• Proficient understanding of HPC and AI solution technologies, including CPUs, GPUs, high-speed interconnects, and supporting software.

• Experience with job scheduling workloads and orchestration tools like Slurm and Kubernetes.

• Strong knowledge of Windows and Linux operating systems, particularly Red Hat/CentOS and Ubuntu.

• Familiarity with networking concepts, sockets, firewalld, iptables, Wireshark, ACLs, OS-level security measures, TCP, DHCP, and DNS.

• Experience with storage solutions including Lustre, GPFS, and Weka.io.

• Proficient in Python programming and Bash scripting.

• Familiarity with automation tools such as Jenkins, Ansible, Puppet, or Chef.

• In-depth knowledge of InfiniBand and Ethernet networking protocols.

• Experience with virtual systems like VMware, Hyper-V, KVM, or Citrix.

• Familiarity with cloud platforms such as AWS, Azure, or Google Cloud.

• Preferred: knowledge of CPU and/or GPU architecture.

• Preferred: understanding of Kubernetes and container-related microservice technologies.

• Preferred: experience with GPU-oriented hardware/software like DGX and CUDA.

• Preferred: experience with RDMA fabrics, including InfiniBand or RoCE.


🏝️ Benefits

• Equal opportunity employer.

• Reasonable accommodations for individuals with disabilities to engage in the application or interview process, perform essential job functions, and receive other employment-related benefits and privileges.

People also viewed

Mercor5 hours ago

AI Safety Red Teamer

US flagUnited States OnlyFreelanceArtificial Intelligence$70 – $84/hour
ApplyView job
Mercor6 hours ago

AI Safety Expert – English, Telugu

US flagUnited States OnlyFreelanceArtificial Intelligence$16 – $22/hour
ApplyView job
Mercor6 hours ago

AI Safety Experts, English, Punjabi

US flagUnited States OnlyFreelanceArtificial Intelligence$16 – $22/hour
ApplyView job
Mercor15 hours ago

AI Safety Expert, English, Gujarati

US flagUnited States OnlyFreelanceArtificial Intelligence$16 – $22/hour
ApplyView job
Mercor15 hours ago

AI Safety Experts – English, Punjabi

US flagUnited States OnlyFreelanceArtificial Intelligence$16 – $22/hour
ApplyView job
Mercor15 hours ago

AI Safety Expert – English, Gujarati

US flagUnited States OnlyFreelanceArtificial Intelligence$16 – $22/hour
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers