Remotery

Research Engineer, Reinforcement Learning

atLiveKitRemoteUS flagUnited StatesFull-timeResearch EngineerMid-levelSenior$135k – $300k/year

Posted 1 day ago

This is a fully remote position, open to applicants in United States.

📋 Description

• Develop the environments and models for verifiers to train against.

• Take ownership of the synthetic data pipeline from its generation through to quality assurance gates.

• Execute comprehensive training experiments and clarify modifications made to the model.

• Create evaluations that releases must successfully complete.

• Select and modify open-weight base models tailored for LiveKit’s objectives.

• Ensure that the trained behavior is effective for both voice and text agents.

• Deploy models into production and continually enhance them based on actual usage data.

• Establish post-training functionalities for agents operating across voice, SMS, and chat platforms, including managing long-horizon context and dependable tool usage.


⛳️ Requirements

• Proficient Python engineering capabilities.

• Proven experience in guiding a model from raw data to production.

• Ability to approach data as a product, focusing on aspects such as coverage, diversity, and leakage.

• Skill in designing against models that exploit weak rewards.

• Comfortable working with GPUs and knowledgeable about their constraints.

• Ability to discern when training should occur and when it should not.

• Ability to collaborate effectively in a remote work setting.

• Experience in post-training processes, fine-tuning, reward design, or reinforcement learning methodologies like GRPO.

• Familiarity with RL and fine-tuning frameworks such as TRL, verl, OpenRLHF, or custom training loops.

• Background in rapid rollouts utilizing vLLM or SGLang.

• Experience in multi-GPU training using FSDP.

• Expertise in training tool-using or multi-turn agents.

• Experience with execution sandboxes, verifiers, evaluation harnesses, or similar tools.

• Knowledge of open-weight model families like Qwen or Llama.

• Proficiency with LoRA or comparable techniques.

• Legal authorization to work in the US or the country of residence.

• Must specify if future sponsorship support is needed.


🏝️ Benefits

• Equity package.

• Health benefits.

• Dental benefits.

• Vision benefits.

• Flexible vacation policy.

• Opportunity to influence the brand of a rapidly growing developer platform.

• Collaboration with a small, senior team that values craftsmanship and creativity.

People also viewed

NetflixAug 13

Research Engineer 4/5 – Member Lifecycle and Monetization

US flagUnited States OnlyFull-timeResearch Engineer$466k – $750k/year
ApplyView job
XBOXAug 13

Machine Learning Research Engineer – Central Technology

US flagCalifornia OnlyFull-timeResearch Engineer$79.2k – $146.5k/year
ApplyView job
Wells FargoAug 11

Senior Lead Application Pen Testing, Cybersecurity Research Engineer

US flagArizona, +3 more statesFull-timeResearch Engineer$159k – $305k/year
ApplyView job
AvengaAug 11

Senior Operations Research Engineer

PL flagPoland OnlyFull-timeResearch Engineer
ApplyView job
BuiltAug 8

Senior Innovation Engineer, Marketplace

US flagTennessee OnlyFull-timeResearch Engineer$140k – $180k/year
ApplyView job
Weekday (YC W21)Aug 7

Research Engineer

IN flagIndia OnlyFull-timeResearch Engineer₹1.5M – ₹2.5M/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers