
Research Scientist / Engineer β Training Infrastructure
Posted 2 days ago

Posted 2 days ago
This is a fully remote position, open to applicants in Europe.
β’ Design, develop, and enhance efficient distributed training systems for models utilizing thousands of GPUs.
β’ Investigate and implement cutting-edge parallelization techniques, such as FSDP, Tensor Parallel, Pipeline Parallel, and Expert Parallel.
β’ Create monitoring, visualization, and debugging tools for extensive training processes.
β’ Enhance training stability, convergence, and resource efficiency across large clusters.
β’ Familiarize oneself with the existing training stack and troubleshoot stability and utilization challenges on a large scale.
β’ Achieve a parallelization or stability enhancement that significantly impacts a real training run.
β’ Develop monitoring and tools to ensure large training runs are reliable and efficient.
β’ Extensive experience in distributed PyTorch training and parallelism within foundation-model training.
β’ Profound knowledge of GPU clusters, networking, and storage systems.
β’ Familiarity with communication libraries (NCCL, MPI) and optimization of distributed systems.
β’ Strong skills in Linux systems administration and scripting (preferred).
β’ Experience managing training runs with over 100 GPUs (preferred).
β’ Knowledge of containerization, orchestration, and cloud infrastructure (preferred).
β’ Proficiency in FSDP and multi-node training.
β’ Capability to work remotely within the EU.
β’ Equal opportunity employer.
β’ Participation in a voluntary diversity and inclusion survey; opting out will not impact the job application.
β’ Remote work flexibility.
Praxis
Praxis
DECA
Get handpicked remote jobs straight to your inbox weekly.