
Senior Deep Learning Software Infrastructure Engineer
Posted Aug 7

Posted Aug 7
This is a fully remote position, open to applicants in California.
• Design, enhance, and solidify deep learning infrastructure libraries and frameworks for training on clusters with thousands of GPUs.
• Optimize the training stack's efficiency, which encompasses data loaders, distributed training, scheduling, and performance monitoring.
• Develop reliable training pipelines and libraries for extensive video datasets and facilitate rapid experimentation.
• Collaborate with researchers, model engineers, and internal platform teams to boost efficiency, minimize interruptions, and enhance training availability.
• Take ownership of essential infrastructure components, including orchestration libraries, distributed training frameworks, and fault-tolerant training systems.
• Work alongside leadership to scale infrastructure in accordance with increasing GPU capacity and dataset size while ensuring developer efficiency and stability.
• Bachelor’s, Master’s, or PhD in Computer Science, Electrical/Computer Engineering, or a related discipline, or equivalent experience.
• Over 12 years of professional experience in developing and scaling high-performance distributed systems, preferably in ML, HPC, or large-scale data infrastructure.
• In-depth understanding of deep learning frameworks, particularly PyTorch.
• Familiarity with large-scale training techniques, including DDP/FSDP, NCCL, tensor parallelism, and pipeline parallelism.
• Experience in performance profiling.
• Strong background in datacenter networking systems, such as RoCE and IB.
• Experience working with parallel filesystems, including Lustre.
• Knowledge of storage systems and schedulers like Slurm and Kubernetes.
• Proficient in Python, with experience in developing production-grade libraries, orchestration layers, and automation tools.
• Capable of collaborating with ML researchers, infrastructure engineers, and product leads to translate requirements into effective systems.
• Experience in scaling GPU training clusters exceeding 1,000 GPUs.
• Expertise in fault tolerance and high availability, including elastic training and large-scale observability.
• Proven hands-on technical leadership, with the ability to set guidelines for ML systems engineering.
• Equity
• Benefits
adconova GmbH
Teleperformance
Trilon Group
Carbon60
Get handpicked remote jobs straight to your inbox weekly.