Software Engineer, Golang, Slurm

Posted Aug 22

This is a fully remote position, open to applicants in Cyprus, +3 more countries.

📋 Description

• Design and develop a managed Slurm service on a Kubernetes platform.

• Produce clean, dependable, and maintainable code in Go.

• Create scheduling and orchestration functionalities tailored for GPU-intensive and distributed workloads.

• Establish observability and automated remediation for failures related to GPU, nodes, networks, and control planes, utilizing VictoriaMetrics, Grafana, and DCGM.

• Maintain traditional Slurm cluster functionalities while operating infrastructure on Kubernetes.

• Identify performance and reliability challenges across GPUs, schedulers, hardware, high-performance networks, and distributed storage solutions.

• Assume end-to-end responsibility for complex challenges within distributed systems.

• Collaborate with technology partners and a global team dedicated to building infrastructure and software for AI, cloud, networking, and security.


⛳️ Requirements

• Practical experience using Slurm in a production environment from a user’s standpoint, including submitting and debugging workloads via sbatch, srun, squeue, and sinfo.

• Strong expertise in Go programming.

• Experience in creating production-grade Kubernetes operators, controllers, CRDs, and reconciliation loops.

• Knowledge of maintaining traditional Slurm cluster functionalities while running the base infrastructure on Kubernetes.

• Proven experience in diagnosing performance and reliability issues across GPUs, schedulers, hardware, high-performance networks, and distributed storage systems.

• A product-oriented mindset with robust customer empathy.

• Exceptional communication skills and the ability to take comprehensive ownership of complex distributed system challenges.

• Nice to have: experience managing large-scale HPC or GPU clusters for external clients.

• Nice to have: familiarity with PyTorch distributed training and other large-scale AI/ML frameworks.

• Nice to have: knowledge of InfiniBand, RoCE, RDMA, GPUDirect, Lustre, WEKA, Ceph, or similar high-performance infrastructures.

• Nice to have: experience in building unified job-submission workflows across Kubernetes and Slurm.

• Nice to have: experience in GPU-cloud or HPC product engineering environments.

• Nice to have: contributions to Slurm, Kubernetes, Soperator, or other cloud-native and HPC open-source initiatives.


🏝️ Benefits

• Competitive salary package.

• Flexible working hours.

• Options for hybrid or remote work, depending on your position.

• Opportunity to work from any location globally for up to 45 days annually.

• Comprehensive private medical insurance for you and your family.*

• Additional paid vacation and sick leave days.*

• Support for significant life events and celebrations.

• Language learning opportunities.

• Modern and inviting offices stocked with snacks, beverages, and entertainment.*

• Team sports and social engagement activities.*

People also viewed

Arista Networks1 day ago

Software Engineer, Kernel and BIOS

HU flagHungary OnlyFull-timeFull-stack Engineer
ApplyView job
Coforma1 day ago

Full Stack Engineer

US flagArizona, +19 more statesFull-timeFull-stack Engineer$119.8k – $149.2k/year
ApplyView job
Platform.sh1 day ago

Senior Software Engineer

ES flagSpain, +5 more countriesFull-timeFull-stack Engineer
ApplyView job
Arista Networks1 day ago

Senior Software Engineer – Layer1, C++

IE flagIreland OnlyFull-timeFull-stack Engineer
ApplyView job
Wikimedia Foundation1 day ago

Senior Software Engineer, iOS

US flagArizona, +31 more statesFull-timeFull-stack Engineer$56 – $87/hour
ApplyView job
Wikimedia Foundation1 day ago

Senior Software Engineer, Core Experiences – Contract

US flagArizona, +31 more statesFull-timeFull-stack Engineer$56 – $87/hour
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers