Remotery

Senior HPC AI Cluster Engineer

atNVIDIARemoteUS flagCaliforniaFull-timeArtificial IntelligenceSenior$176k – $333.5k/year

Posted 18 hours ago

This is a fully remote position, open to applicants in California.

📋 Description

• Design, implement, and support extensive HPC/AI clusters, ensuring effective monitoring, logging, and alerting systems are in place.

• Oversee Linux job and workload scheduling alongside orchestration tools.

• Create and maintain continuous integration and delivery pipelines.

• Develop tools to automate the deployment and management of large-scale infrastructure environments.

• Automate operational monitoring and alerting, facilitating self-service resource consumption.

• Install monitoring solutions for servers, networks, and storage systems.

• Troubleshoot issues from bare metal through the operating system, software stack, and application level.

• Develop, refine, and document standard methodologies for internal teams.

• Provide support for Research & Development activities and engage in POCs/POVs aimed at future enhancements.

• Collaborate with HPC, OS, GPU compute, and systems specialists to design, develop, and launch large-scale performance platforms.

• Share insights on system design and tuning strategies for large-scale computing tasks.

• Work alongside accelerated computing and deep learning software and hardware platforms, researchers, developers, and customers to enhance workflows and create unique solutions.


⛳️ Requirements

• A degree in Computer Science, Engineering, or a related discipline (or equivalent experience).

• Over 8 years of relevant experience.

• Familiarity with HPC and AI solution technologies, including CPUs, GPUs, high-speed interconnects, and supporting software.

• Proficiency in job scheduling workloads and orchestration tools such as Slurm and Kubernetes.

• Strong knowledge of Windows and Linux (Red Hat/CentOS and Ubuntu) networking and internals, ACLs, OS-level security, and common protocols like TCP, DHCP, and DNS.

• Experience with storage solutions including Lustre, GPFS, and Weka.io.

• Awareness of emerging storage technologies.

• Proficient in Python programming and Bash scripting.

• Experience with automation and configuration management tools such as Jenkins, Ansible, Puppet, and Chef.

• In-depth knowledge of networking protocols like InfiniBand and Ethernet.

• Strong understanding and experience with virtual systems such as VMware, Hyper-V, KVM, or Citrix.

• Familiarity with cloud computing platforms such as AWS, Azure, and Google Cloud.

• Knowledge of CPU and/or GPU architecture.

• Familiarity with Kubernetes and container-related microservice technologies.

• Experience with GPU-centric hardware/software like DGX and CUDA.

• Experience with RDMA (InfiniBand or RoCE) fabrics.


🏝️ Benefits

• Competitive salaries.

• Generous benefits package.

• Equity.

People also viewed

CVS Health9 hours ago

Lead Director, Digital Product – Conversational AI Strategy

US flagTexas OnlyFull-timeArtificial Intelligence$144.2k – $288.4k/year
ApplyView job
One Impression9 hours ago

AI Generalist Intern – Founder's Office

IN flagIndia OnlyInternshipArtificial Intelligence
ApplyView job
Volga Partners10 hours ago

AI Language Quality Evaluator – Greek/English, Mid-Level

GR flagGreece, +1 more countryFreelanceArtificial Intelligence$7 – $9/hour
ApplyView job
Mercor10 hours ago

Senior Design Expert – Paid AI Design Research Study

US flagUnited States OnlyFreelanceArtificial Intelligence$150 – $250/hour
ApplyView job
Reveleer10 hours ago

SVP, AI & Data

US flagUnited States OnlyFull-timeArtificial Intelligence$305k – $355k/year
ApplyView job
Mitratech11 hours ago

AI Automation Specialist

MX flagMexico OnlyFull-timeArtificial Intelligence
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers