Remotery

Senior Software Engineer, RL Post-Training Frameworks

Posted Jul 3

This is a fully remote position, open to applicants in Germany.

πŸ“‹ Description

β€’ Design and construct RL post-training infrastructure that efficiently scales from experimentation on a single GPU to deployment across thousands of nodes.

β€’ Optimize RL training-inference-rollout loops on GPUs, CPUs, and LPUs for critical performance enhancements.

β€’ Contribute to and enhance the performance and usability of open-source RL frameworks.

β€’ Collaborate with teams developing CPU-driven rollout workloads, encompassing tool usage, code execution, and agentic environments.

β€’ Represent the needs of researchers and partners with NVIDIA's networking, math library, and compiler teams.


⛳️ Requirements

β€’ MS or PhD in Computer Science, Computer Engineering, or a related discipline (or equivalent experience).

β€’ Over 5 years of professional experience in distributed systems, high-performance computing, deep learning infrastructure, or ML systems engineering.

β€’ Strong expertise in Python and C/C++.

β€’ Proven experience in building or contributing to large-scale distributed systems or runtime frameworks in production at a leading AI lab, hyperscaler, or major tech company.

β€’ Excellent verbal and written communication skills with the ability to collaborate across organizational and geographic boundaries.

β€’ In-depth knowledge in one or more of the following technical areas: Reinforcement learning for LLM post-training (RLHF, PPO, GRPO, DPO, reward modeling), including the mapping of algorithms to distributed execution and the related system challenges (heterogeneous placement, rollouts, environment execution, resharding between training and generation).

β€’ Familiarity with PyTorch internals, including distributed training primitives (FSDP, tensor parallelism, pipeline parallelism) and their integration.

β€’ Understanding of Kubernetes runtime internals (container lifecycle, pod scheduling, resource quotas, GPU allocation).

β€’ Comprehensive knowledge of end-to-end distributed systems design (service boundaries, data flows, consistency models, failure modes, recovery strategies).


🏝️ Benefits

β€’ Highly competitive salaries.

β€’ Comprehensive benefits package.

People also viewed

Future Connections2 days ago

Senior Fullstack Engineer

ES flagSpain OnlyFull-timeFull-stack Engineer
ApplyView job
Akamai Technologies2 days ago

Software Engineer II – Data Products

PL flagPoland OnlyFull-timeFull-stack Engineer
ApplyView job
cigus2 days ago

Senior PLC Software Developer, CODESYS V3

DE flagGermany OnlyFull-timeFull-stack Engineer
ApplyView job
Digi International2 days ago

Principal GTM Engineer

US flagUnited States OnlyFull-timeFull-stack Engineer
ApplyView job
Fivetran2 days ago

Staff Software Engineer, Metadata

US flagCalifornia OnlyFull-timeFull-stack Engineer$183k – $245k/year
ApplyView job
CQ fluency2 days ago

Full-Stack DevSecOps Engineer

PH flagPhilippines, +2 more statesFull-timeFull-stack Engineer
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers