
Research Engineer β Post-Training
Posted 3 days ago

Posted 3 days ago
This is a fully remote position, open to applicants in United States, +1 more country.
β’ Develop the end-to-end RL post-training infrastructure, encompassing rollout ingestion, reward calculation, policy modifications, and the distribution of updated weights throughout the network.
β’ Establish the vision for the post-training stack and oversee its development.
β’ Modify RL algorithms to function effectively in asynchronous, high-latency, and partially trusted generation environments.
β’ Tackle issues related to staleness tolerance, off-policy adjustments, and communication-efficient policy enhancements.
β’ Create evaluations that showcase improvements in model performance.
β’ Release the first decentralized post-trained model as a publicly accessible artifact.
β’ Practical experience in executing RL post-training on large language models, such as RLHF, RLVR, or reasoning-oriented RL.
β’ Familiarity with rollout generation, asynchronous training processes, and weight synchronization.
β’ Proficient skills in Python and PyTorch at a production level.
β’ Experience with concurrency, failure management, and profiling prior to optimization.
β’ Publications in areas related to RL post-training, asynchronous or distributed RL, or similar fields, or unpublished work that can be thoroughly defended.
β’ A strong belief in Protocol Learning as a feasible approach for collective, trustless, and sovereign AI.
β’ Proficient in English, both written and spoken, at a professional level.
β’ Ability to collaborate effectively across different time zones.
β’ Nice to have: experience in training over slow networks or with decentralized/federated configurations.
β’ Nice to have: knowledge of vLLM or SGLang serving internals.
β’ Nice to have: experience with reward modeling or datasets that verify rewards.
β’ Nice to have: familiarity with P2P networking and NAT traversal.
β’ Nice to have: experience in proprietary, open-weight, and open-source AI laboratories.
β’ Significant equity ownership for key technical contributors, in addition to a competitive base salary.
β’ Flexible working environment with team members located globally.
β’ Optional full visa sponsorship and relocation assistance to either Australia or the US.
InspiredOne
AssemblyAI
Pluralis Research
Pluralis Research
Get handpicked remote jobs straight to your inbox weekly.