
Research Scientist / Engineer β Reinforcement Learning Infrastructure
Posted 2 days ago

Posted 2 days ago
This is a fully remote position, open to applicants in Europe.
β’ Design, develop, and enhance distributed reinforcement learning post-training systems utilizing thousands of GPUs.
β’ Create high-throughput rollout generation systems that incorporate vLLM, SGLang, weight synchronization, and asynchronous/off-policy strategies.
β’ Engineer scalable reinforcement learning environments for agentic, multi-step tasks, including sandboxed code execution, tool usage, computer interaction, and multimodal interfaces.
β’ Construct reward infrastructure featuring verifiable/programmatic rewards, reward-model serving, LLM-as-judge pipelines, and safeguards against reward manipulation.
β’ Create evaluation, monitoring, and debugging tools to ensure stable large-scale reinforcement learning operations.
β’ Enhance training efficiency and reliability while transforming post-training concepts into production implementations alongside researchers.
β’ Acquire knowledge of the existing RL stack, identify bottlenecks, implement and validate enhancements, and reinforce the complete loop across thousands of GPUs.
β’ Practical experience with post-training language models leveraging reinforcement learning (PPO/GRPO-family, RLHF, RLVR) at significant scale.
β’ In-depth experience with distributed PyTorch training and parallelism (FSDP, Tensor/Pipeline/Expert Parallel) aimed at foundation models.
β’ Proven track record in constructing reinforcement learning environments, reward functions, verifiers, or evaluation harnesses for language model agents, including sandboxed execution and multi-turn tool use.
β’ Profound understanding of RL post-training frameworks (veRL, OpenRLHF, TRL, Ray orchestration) and rollout inference engines (vLLM, SGLang).
β’ Strong grasp of GPU clusters, networking, and communication libraries (NCCL, MPI) under mixed training and inference scenarios.
β’ Experience with containerization and orchestration (Kubernetes, Ray) for large fleets of environments and sandboxed workloads.
β’ Contributions to research in reinforcement learning for language models, or open-source contributions to RL training frameworks.
β’ Equal opportunity employer.
β’ Flexible remote work arrangements available within the EU.
Praxis
Praxis
DECA
Get handpicked remote jobs straight to your inbox weekly.