
Applied Research Scientist, LLM Evaluation – Post-Training
Posted Jul 27

Posted Jul 27
This is a fully remote position, open to applicants in United States.
• Lead research and experimentation to understand how evaluation design, measurement strategies, and feedback signals impact model enhancement.
• Assist in shaping the future of evaluation-driven model improvement workflows.
• Investigate how various evaluation methods (human, automated, hybrid) influence model selection and outcomes after training.
• Create experiments that yield credible and actionable insights.
• Support customer interactions by applying scientific rigor to evaluation strategies, methodology assessments, and technical guidance.
• Establish and implement a research agenda centered on LLM evaluation and post-training, particularly in evaluation-driven model enhancement.
• Design robust experiments to examine the effects of evaluation methodologies on fine-tuning and outcomes following training.
• Develop and validate evaluation frameworks for LLM and multimodal systems.
• Analyze model behavior and identify failure patterns; provide actionable recommendations for model enhancement and evaluation redesign.
• Collaborate with AI/ML Research Engineers to transform research methods into scalable evaluation and post-training workflows.
• MS/PhD in Computer Science, Machine Learning, Statistics, Applied Mathematics, AI, or a related quantitative scientific discipline (PhD is strongly preferred).
• Over 5 years of pertinent experience in applied research or research science in ML/AI, particularly significant work with LLMs or foundation models.
• Proven experience in LLM evaluation, benchmarking, alignment, post-training, or model quality research.
• Strong background in experimental design, statistical analysis, and scientific reasoning in ML systems.
• Proficient coding skills in Python for research experimentation and analysis (e.g., data processing, evaluation pipelines, statistical analysis, visualization).
• Experience with contemporary ML tools/frameworks (e.g., PyTorch, Hugging Face, JAX/TensorFlow as applicable) sufficient for designing and conducting model/evaluation experiments.
• Capability to assess and compare human and automated evaluation methods, including trade-offs in cost, reliability, validity, and scalability.
• Experience in designing evaluation studies and protocols that are reproducible across datasets, model versions, and evaluation runs.
• Ability to collaborate directly with technical stakeholders, including research scientists, ML engineers, data scientists, and technical representatives from customers.
• Excellent communication skills and the ability to clearly present nuanced technical findings, assumptions, and limitations.
• Competitive salary and performance-based bonuses.
• Comprehensive health, dental, and vision insurance.
• Flexible work hours and remote work options.
• Opportunities for professional development and continuous learning.
• Collaborative and innovative work environment.
University of Arkansas System
WashU IT
HMH
Arizona
Get handpicked remote jobs straight to your inbox weekly.