Remotery

Senior Deep Learning Software Infrastructure Engineer

atNVIDIARemoteUS flagCaliforniaFull-timeInfrastructure EngineerSenior$224k – $431.3k/year

Posted Aug 7

This is a fully remote position, open to applicants in California.

📋 Description

• Design, enhance, and solidify deep learning infrastructure libraries and frameworks for training on clusters with thousands of GPUs.

• Optimize the training stack's efficiency, which encompasses data loaders, distributed training, scheduling, and performance monitoring.

• Develop reliable training pipelines and libraries for extensive video datasets and facilitate rapid experimentation.

• Collaborate with researchers, model engineers, and internal platform teams to boost efficiency, minimize interruptions, and enhance training availability.

• Take ownership of essential infrastructure components, including orchestration libraries, distributed training frameworks, and fault-tolerant training systems.

• Work alongside leadership to scale infrastructure in accordance with increasing GPU capacity and dataset size while ensuring developer efficiency and stability.


⛳️ Requirements

• Bachelor’s, Master’s, or PhD in Computer Science, Electrical/Computer Engineering, or a related discipline, or equivalent experience.

• Over 12 years of professional experience in developing and scaling high-performance distributed systems, preferably in ML, HPC, or large-scale data infrastructure.

• In-depth understanding of deep learning frameworks, particularly PyTorch.

• Familiarity with large-scale training techniques, including DDP/FSDP, NCCL, tensor parallelism, and pipeline parallelism.

• Experience in performance profiling.

• Strong background in datacenter networking systems, such as RoCE and IB.

• Experience working with parallel filesystems, including Lustre.

• Knowledge of storage systems and schedulers like Slurm and Kubernetes.

• Proficient in Python, with experience in developing production-grade libraries, orchestration layers, and automation tools.

• Capable of collaborating with ML researchers, infrastructure engineers, and product leads to translate requirements into effective systems.

• Experience in scaling GPU training clusters exceeding 1,000 GPUs.

• Expertise in fault tolerance and high availability, including elastic training and large-scale observability.

• Proven hands-on technical leadership, with the ability to set guidelines for ML systems engineering.


🏝️ Benefits

• Equity

• Benefits

People also viewed

adconova GmbH19 hours ago

Senior Cloud & AI Infrastructure Engineer

DE flagGermany OnlyFull-timeInfrastructure Engineer€60k – €100k/year
ApplyView job
Teleperformance1 day ago

Senior Systems Engineer – IT Infrastructure Engineer

HR flagCroatia OnlyFull-timeInfrastructure Engineer
ApplyView job
Trilon Group2 days ago

AWS Infrastructure Engineer

US flagUnited States OnlyFull-timeInfrastructure Engineer$90k – $110k/year
ApplyView job
Carbon602 days ago

Principal Infrastructure Architect

CA flagCanada OnlyFull-timeInfrastructure EngineerC$180k – C$220k/year
ApplyView job
fal2 days ago

Senior/Staff Kubernetes Infrastructure Engineer

US flagUnited States OnlyFull-timeInfrastructure Engineer$180k – $250k/year
ApplyView job
LatamCent2 days ago

Lead Security and Infrastructure Engineer

US flagFlorida OnlyFull-timeInfrastructure Engineer$170k – $210k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers