Remotery

Senior HPC AI Cluster Engineer

Posted 3 days ago

This is a fully remote position, open to applicants in Switzerland.

📋 Description

• Design, implement, and maintain extensive HPC/AI clusters, including monitoring, logging, and alerting systems.

• Oversee Linux job/workload schedules and orchestration tools.

• Create and sustain continuous integration and delivery pipelines.

• Develop tools to automate the deployment and management of large-scale infrastructure environments, including operational monitoring and alerting, as well as self-service resource usage.

• Implement monitoring solutions for servers, networks, and storage systems.

• Diagnose issues from bare metal through to the operating system, software stack, and application levels.

• Create, refine, and document standard practices for internal teams.

• Assist Research & Development initiatives and participate in POCs/POVs for future enhancements.

• Offer insights on system design at scale and tuning strategies for large-scale computational tasks.

• Collaborate with HPC, OS, GPU compute, and systems specialists to design, develop, and establish large-scale performance platforms.

• Engage with accelerated computing and deep learning software and hardware platforms, researchers, developers, and clients to enhance workflows and create unique solutions.


⛳️ Requirements

• A degree in Computer Science, Engineering, or a related discipline.

• Over 8 years of relevant experience.

• Proficiency in HPC and AI solution technologies involving CPUs, GPUs, high-speed interconnects, and related software.

• Familiarity with workload scheduling and orchestration tools like Slurm and Kubernetes.

• Strong command of Windows and Linux operating systems, including Red Hat/CentOS and Ubuntu.

• Understanding of networking concepts, including sockets, firewalld, iptables, Wireshark, networking internals, ACLs, OS-level security measures, TCP, DHCP, DNS, and common protocols.

• Experience with storage solutions such as Lustre, GPFS, and Weka.io.

• Proficient in Python programming and Bash scripting.

• Familiarity with automation and configuration management tools such as Jenkins, Ansible, Puppet, and Chef.

• In-depth knowledge of networking protocols, including InfiniBand and Ethernet.

• Comprehensive understanding and experience with virtual systems like VMware, Hyper-V, KVM, or Citrix.

• Acquainted with cloud computing platforms, including AWS, Azure, and Google Cloud.

• Knowledge of CPU and/or GPU architectures.

• Familiarity with Kubernetes and microservice technologies related to containers.

• Experience with GPU-centric hardware/software such as DGX and CUDA.

• Knowledge of RDMA fabrics like InfiniBand or RoCE.


🏝️ Benefits

• Competitive salary and performance-based incentives.

• Comprehensive health and wellness benefits.

• Opportunities for professional development and continuous learning.

• Flexible work arrangements and a supportive work environment.

People also viewed

Mashreq10 hours ago

VP, AI Audit – Shared Services

IN flagIndia OnlyFull-timeArtificial Intelligence
ApplyView job
LILT AI11 hours ago

AI Training Contributor, Luxembourgish

LU flagLuxembourg OnlyFreelanceArtificial Intelligence
ApplyView job
BCD Travel11 hours ago

Head of AI

GB flagUnited Kingdom, +2 more statesFull-timeArtificial Intelligence€160k – €200k/year
ApplyView job
BCD Travel11 hours ago

AI Business Transformation Lead

GB flagUnited Kingdom, +2 more statesFull-timeArtificial Intelligence€110k – €190k/year
ApplyView job
One Impression14 hours ago

AI Generalist Intern

IN flagIndia OnlyInternshipArtificial Intelligence
ApplyView job
Creative Chaos15 hours ago

Principal AI Process Orchestrator

BR flagBrazil, +5 more statesFull-timeArtificial Intelligence
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers