
Senior HPC AI Cluster Engineer
Posted 3 days ago

Posted 3 days ago
This is a fully remote position, open to applicants in Switzerland.
• Design, implement, and maintain extensive HPC/AI clusters, including monitoring, logging, and alerting systems.
• Oversee Linux job/workload schedules and orchestration tools.
• Create and sustain continuous integration and delivery pipelines.
• Develop tools to automate the deployment and management of large-scale infrastructure environments, including operational monitoring and alerting, as well as self-service resource usage.
• Implement monitoring solutions for servers, networks, and storage systems.
• Diagnose issues from bare metal through to the operating system, software stack, and application levels.
• Create, refine, and document standard practices for internal teams.
• Assist Research & Development initiatives and participate in POCs/POVs for future enhancements.
• Offer insights on system design at scale and tuning strategies for large-scale computational tasks.
• Collaborate with HPC, OS, GPU compute, and systems specialists to design, develop, and establish large-scale performance platforms.
• Engage with accelerated computing and deep learning software and hardware platforms, researchers, developers, and clients to enhance workflows and create unique solutions.
• A degree in Computer Science, Engineering, or a related discipline.
• Over 8 years of relevant experience.
• Proficiency in HPC and AI solution technologies involving CPUs, GPUs, high-speed interconnects, and related software.
• Familiarity with workload scheduling and orchestration tools like Slurm and Kubernetes.
• Strong command of Windows and Linux operating systems, including Red Hat/CentOS and Ubuntu.
• Understanding of networking concepts, including sockets, firewalld, iptables, Wireshark, networking internals, ACLs, OS-level security measures, TCP, DHCP, DNS, and common protocols.
• Experience with storage solutions such as Lustre, GPFS, and Weka.io.
• Proficient in Python programming and Bash scripting.
• Familiarity with automation and configuration management tools such as Jenkins, Ansible, Puppet, and Chef.
• In-depth knowledge of networking protocols, including InfiniBand and Ethernet.
• Comprehensive understanding and experience with virtual systems like VMware, Hyper-V, KVM, or Citrix.
• Acquainted with cloud computing platforms, including AWS, Azure, and Google Cloud.
• Knowledge of CPU and/or GPU architectures.
• Familiarity with Kubernetes and microservice technologies related to containers.
• Experience with GPU-centric hardware/software such as DGX and CUDA.
• Knowledge of RDMA fabrics like InfiniBand or RoCE.
• Competitive salary and performance-based incentives.
• Comprehensive health and wellness benefits.
• Opportunities for professional development and continuous learning.
• Flexible work arrangements and a supportive work environment.
Mashreq
LILT AI
BCD Travel
BCD Travel
Get handpicked remote jobs straight to your inbox weekly.