Senior Deep Learning Software Infrastructure Engineer

atNVIDIARemoteUS flagCaliforniaFull-timeInfrastructure EngineerSenior$224k – $431.3k/year

Posted Aug 7

This is a fully remote position, open to applicants in California.

📋 Description

• Design, enhance, and solidify deep learning infrastructure libraries and frameworks for training on clusters with thousands of GPUs.

• Optimize the training stack's efficiency, which encompasses data loaders, distributed training, scheduling, and performance monitoring.

• Develop reliable training pipelines and libraries for extensive video datasets and facilitate rapid experimentation.

• Collaborate with researchers, model engineers, and internal platform teams to boost efficiency, minimize interruptions, and enhance training availability.

• Take ownership of essential infrastructure components, including orchestration libraries, distributed training frameworks, and fault-tolerant training systems.

• Work alongside leadership to scale infrastructure in accordance with increasing GPU capacity and dataset size while ensuring developer efficiency and stability.


⛳️ Requirements

• Bachelor’s, Master’s, or PhD in Computer Science, Electrical/Computer Engineering, or a related discipline, or equivalent experience.

• Over 12 years of professional experience in developing and scaling high-performance distributed systems, preferably in ML, HPC, or large-scale data infrastructure.

• In-depth understanding of deep learning frameworks, particularly PyTorch.

• Familiarity with large-scale training techniques, including DDP/FSDP, NCCL, tensor parallelism, and pipeline parallelism.

• Experience in performance profiling.

• Strong background in datacenter networking systems, such as RoCE and IB.

• Experience working with parallel filesystems, including Lustre.

• Knowledge of storage systems and schedulers like Slurm and Kubernetes.

• Proficient in Python, with experience in developing production-grade libraries, orchestration layers, and automation tools.

• Capable of collaborating with ML researchers, infrastructure engineers, and product leads to translate requirements into effective systems.

• Experience in scaling GPU training clusters exceeding 1,000 GPUs.

• Expertise in fault tolerance and high availability, including elastic training and large-scale observability.

• Proven hands-on technical leadership, with the ability to set guidelines for ML systems engineering.


🏝️ Benefits

• Equity

• Benefits

People also viewed

Second Nature2 days ago

Cloud Specialist – Infrastructure Senior

AR flagArgentina OnlyFull-timeInfrastructure Engineer
ApplyView job
Headway2 days ago

Senior Software Engineer – Infrastructure

US flagUnited States OnlyFull-timeInfrastructure Engineer$182.8k – $278.5k/year
ApplyView job
SYNCREON3 days ago

Senior Mainframe Systems Programmer – Infrastructure Engineering, SMP/E

US flagNew Jersey OnlyFull-timeInfrastructure Engineer
ApplyView job
Rentokil Pest Control North America3 days ago

IT Infrastructure Engineer

US flagArizona, +4 more statesFull-timeInfrastructure Engineer$51k – $83.5k/year
ApplyView job
Endava3 days ago

Infrastructure Engineer

FR flagFrance OnlyFull-timeInfrastructure Engineer
ApplyView job
Leidos3 days ago

Cloud Infrastructure Engineer – OCI Engineer

US flagUnited States OnlyFull-timeInfrastructure Engineer$107.9k – $195.1k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers