Research Scientist / Engineer – Reinforcement Learning Infrastructure

Posted 2 days ago

This is a fully remote position, open to applicants in Europe.

πŸ“‹ Description

β€’ Design, develop, and enhance distributed reinforcement learning post-training systems utilizing thousands of GPUs.

β€’ Create high-throughput rollout generation systems that incorporate vLLM, SGLang, weight synchronization, and asynchronous/off-policy strategies.

β€’ Engineer scalable reinforcement learning environments for agentic, multi-step tasks, including sandboxed code execution, tool usage, computer interaction, and multimodal interfaces.

β€’ Construct reward infrastructure featuring verifiable/programmatic rewards, reward-model serving, LLM-as-judge pipelines, and safeguards against reward manipulation.

β€’ Create evaluation, monitoring, and debugging tools to ensure stable large-scale reinforcement learning operations.

β€’ Enhance training efficiency and reliability while transforming post-training concepts into production implementations alongside researchers.

β€’ Acquire knowledge of the existing RL stack, identify bottlenecks, implement and validate enhancements, and reinforce the complete loop across thousands of GPUs.


⛳️ Requirements

β€’ Practical experience with post-training language models leveraging reinforcement learning (PPO/GRPO-family, RLHF, RLVR) at significant scale.

β€’ In-depth experience with distributed PyTorch training and parallelism (FSDP, Tensor/Pipeline/Expert Parallel) aimed at foundation models.

β€’ Proven track record in constructing reinforcement learning environments, reward functions, verifiers, or evaluation harnesses for language model agents, including sandboxed execution and multi-turn tool use.

β€’ Profound understanding of RL post-training frameworks (veRL, OpenRLHF, TRL, Ray orchestration) and rollout inference engines (vLLM, SGLang).

β€’ Strong grasp of GPU clusters, networking, and communication libraries (NCCL, MPI) under mixed training and inference scenarios.

β€’ Experience with containerization and orchestration (Kubernetes, Ray) for large fleets of environments and sandboxed workloads.

β€’ Contributions to research in reinforcement learning for language models, or open-source contributions to RL training frameworks.


🏝️ Benefits

β€’ Equal opportunity employer.

β€’ Flexible remote work arrangements available within the EU.

People also viewed

Praxis2 days ago

Principal Scientist, Process Chemistry

US flagUnited States OnlyFull-timeResearch Scientist$165k – $185k/year
ApplyView job
Praxis2 days ago

Principal Scientist, Formulation Development

US flagUnited States OnlyFull-timeResearch Scientist$170k – $190k/year
ApplyView job
DECA2 days ago

Senior Principal Scientist – US Regulatory Innovation

US flagUnited States OnlyFull-timeResearch Scientist$147k – $191k/year
ApplyView job
Zillow2 days ago

Applied Scientist

DE flagGermany OnlyFull-timeResearch Scientist
ApplyView job
Luma AI2 days ago

Research Scientist / Engineer – Training Infrastructure

EuropeFull-timeResearch Scientist
ApplyView job
Luma AI2 days ago

Research Scientist – Performance Optimization

EuropeFull-timeResearch Scientist
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers