
Senior Software Engineer, RL Post-Training Frameworks
Posted Jul 3

Posted Jul 3
This is a fully remote position, open to applicants in Germany.
β’ Design and construct RL post-training infrastructure that efficiently scales from experimentation on a single GPU to deployment across thousands of nodes.
β’ Optimize RL training-inference-rollout loops on GPUs, CPUs, and LPUs for critical performance enhancements.
β’ Contribute to and enhance the performance and usability of open-source RL frameworks.
β’ Collaborate with teams developing CPU-driven rollout workloads, encompassing tool usage, code execution, and agentic environments.
β’ Represent the needs of researchers and partners with NVIDIA's networking, math library, and compiler teams.
β’ MS or PhD in Computer Science, Computer Engineering, or a related discipline (or equivalent experience).
β’ Over 5 years of professional experience in distributed systems, high-performance computing, deep learning infrastructure, or ML systems engineering.
β’ Strong expertise in Python and C/C++.
β’ Proven experience in building or contributing to large-scale distributed systems or runtime frameworks in production at a leading AI lab, hyperscaler, or major tech company.
β’ Excellent verbal and written communication skills with the ability to collaborate across organizational and geographic boundaries.
β’ In-depth knowledge in one or more of the following technical areas: Reinforcement learning for LLM post-training (RLHF, PPO, GRPO, DPO, reward modeling), including the mapping of algorithms to distributed execution and the related system challenges (heterogeneous placement, rollouts, environment execution, resharding between training and generation).
β’ Familiarity with PyTorch internals, including distributed training primitives (FSDP, tensor parallelism, pipeline parallelism) and their integration.
β’ Understanding of Kubernetes runtime internals (container lifecycle, pod scheduling, resource quotas, GPU allocation).
β’ Comprehensive knowledge of end-to-end distributed systems design (service boundaries, data flows, consistency models, failure modes, recovery strategies).
β’ Highly competitive salaries.
β’ Comprehensive benefits package.
Future Connections
Akamai Technologies
cigus
Digi International
Get handpicked remote jobs straight to your inbox weekly.