
Software Engineer, Golang, Slurm
Posted Aug 22

Posted Aug 22
This is a fully remote position, open to applicants in Cyprus, +3 more countries.
• Design and develop a managed Slurm service on a Kubernetes platform.
• Produce clean, dependable, and maintainable code in Go.
• Create scheduling and orchestration functionalities tailored for GPU-intensive and distributed workloads.
• Establish observability and automated remediation for failures related to GPU, nodes, networks, and control planes, utilizing VictoriaMetrics, Grafana, and DCGM.
• Maintain traditional Slurm cluster functionalities while operating infrastructure on Kubernetes.
• Identify performance and reliability challenges across GPUs, schedulers, hardware, high-performance networks, and distributed storage solutions.
• Assume end-to-end responsibility for complex challenges within distributed systems.
• Collaborate with technology partners and a global team dedicated to building infrastructure and software for AI, cloud, networking, and security.
• Practical experience using Slurm in a production environment from a user’s standpoint, including submitting and debugging workloads via sbatch, srun, squeue, and sinfo.
• Strong expertise in Go programming.
• Experience in creating production-grade Kubernetes operators, controllers, CRDs, and reconciliation loops.
• Knowledge of maintaining traditional Slurm cluster functionalities while running the base infrastructure on Kubernetes.
• Proven experience in diagnosing performance and reliability issues across GPUs, schedulers, hardware, high-performance networks, and distributed storage systems.
• A product-oriented mindset with robust customer empathy.
• Exceptional communication skills and the ability to take comprehensive ownership of complex distributed system challenges.
• Nice to have: experience managing large-scale HPC or GPU clusters for external clients.
• Nice to have: familiarity with PyTorch distributed training and other large-scale AI/ML frameworks.
• Nice to have: knowledge of InfiniBand, RoCE, RDMA, GPUDirect, Lustre, WEKA, Ceph, or similar high-performance infrastructures.
• Nice to have: experience in building unified job-submission workflows across Kubernetes and Slurm.
• Nice to have: experience in GPU-cloud or HPC product engineering environments.
• Nice to have: contributions to Slurm, Kubernetes, Soperator, or other cloud-native and HPC open-source initiatives.
• Competitive salary package.
• Flexible working hours.
• Options for hybrid or remote work, depending on your position.
• Opportunity to work from any location globally for up to 45 days annually.
• Comprehensive private medical insurance for you and your family.*
• Additional paid vacation and sick leave days.*
• Support for significant life events and celebrations.
• Language learning opportunities.
• Modern and inviting offices stocked with snacks, beverages, and entertainment.*
• Team sports and social engagement activities.*
Arista Networks
Coforma
Platform.sh
Arista Networks
Get handpicked remote jobs straight to your inbox weekly.