
Senior Applied Scientist
Posted 16 hours ago

Posted 16 hours ago
This is a fully remote position, open to applicants in United States.
β’ Take charge of LLM post-training processes, encompassing SFT, DPO, and reinforcement fine-tuning (RFT/GRPO), executing training from start to finish on GPU infrastructure.
β’ Develop and train rapid, multi-category reward models that generate scalar signals from contrastive preference pairs.
β’ Create on-policy and real-time assessments by scoring live agent trajectories utilizing LLM-as-a-Judge frameworks.
β’ Convert offline evaluation criteria into adaptable reward functions applicable to various production traces.
β’ Navigate complex and ambiguous tasks from inception to completion in partnership with product, engineering, science, platform, and data teams.
β’ Offer technical guidance and mentorship.
β’ Transform cutting-edge AI research into practical production outcomes.
β’ Master's degree or higher in Computer Science or a related discipline.
β’ Practical experience with LLM post-training and RL fine-tuning: SFT, DPO, RFT/GRPO, conducting end-to-end training on GPU infrastructure.
β’ Familiarity with reward model development; research-level experience is acceptable.
β’ Profound understanding of generative AI, including foundational models, transformers, reinforcement learning, and preference learning.
β’ Capability to independently define and resolve ambiguous challenges from start to finish while providing technical leadership to scientists and MLEs.
β’ Proficient programming abilities, particularly in Python.
β’ Experience with ML frameworks such as PyTorch or TensorFlow.
β’ PhD in Computer Science, Machine Learning, or a related field (preferred).
β’ Background in evaluating agentic AI: multi-turn trace analysis, LLM-as-a-Judge, and understanding the interplay between offline evaluations and online monitoring (preferred).
β’ Published research in post-training, RLHF/RLAIF, preference learning, or reward modeling (preferred).
β’ Familiarity with GPU training platforms such as Databricks or Fireworks (preferred).
β’ Equity awards determined by factors such as experience, performance, and location.
β’ Remote work from a location of your choice.
β’ Flexibility to work from wherever you are most productive.
β’ Support for accommodations related to disabilities or special needs.
AMDEX Corp
Workforce and Community Education
Mercor
Mercor
Get handpicked remote jobs straight to your inbox weekly.