
AI/ML Research Engineer, LLM Post-Training – Evaluation
Posted Jul 27

Posted Jul 27
This is a fully remote position, open to applicants in United States.
• Design and develop the pipelines and tools that link data, evaluation, and post-training processes.
• Assist customers and internal teams in transforming evaluation insights into quantifiable model enhancements.
• Create fine-tuning workflows, such as supervised fine-tuning and preference-based optimization.
• Incorporate evaluation harnesses into the model development lifecycle.
• Enhance the reliability and throughput of experiments.
• Support sophisticated evaluation scenarios, including long-context, cross-modal, and dynamic multi-turn interactions.
• Contribute to Innodata's internal research and development initiatives, including benchmark datasets, evaluation frameworks, and reusable infrastructure for model evaluations and post-training experimentation.
• Lead or co-lead technically intricate ML engineering projects from initial customer discussions to implementation and delivery.
• Design, develop, and enhance LLM training and post-training pipelines, which include data ingestion, preprocessing, fine-tuning, evaluation, and experiment tracking.
• Implement and optimize evaluation systems for LLMs and multimodal models, covering offline benchmarks and task-specific test harnesses.
• Integrate human-in-the-loop and AI-augmented evaluation signals into model development workflows.
• Develop robust infrastructure and tools for reproducible experimentation, metrics logging, and regression monitoring.
• Diagnose model behavior and pipeline failures, addressing issues related to data, training instability, metric discrepancies, and evaluation drift.
• Collaborate with Language Data Scientists and Applied Research Scientists to convert evaluation frameworks into actionable systems.
• Work closely with customer technical stakeholders to comprehend goals, constraints, and success criteria; propose and implement technically sound solutions.
• Contribute to internal research and platform enhancement, including benchmark frameworks, evaluation tools, and improvements to post-training workflows.
• Contribute to best practices and standards for LLM training, evaluation, and quality assurance across various projects.
• Mentor junior engineers and engage in technical design reviews, documentation, and engineering standards across the team.
• BS/MS/PhD in Computer Science, Machine Learning, AI, Applied Mathematics, or a related quantitative technical discipline (MS/PhD preferred).
• 2-3 years of pertinent industry or research engineering experience in ML/AI systems.
• Practical experience with LLM training, fine-tuning, and post-training, including at least one of the following:
• Supervised fine-tuning (SFT).
• Preference optimization (e.g., DPO or similar methods).
• RLHF/RLAIF-style workflows.
• Task- or domain-adaptation of foundation models.
• Strong programming capabilities in Python and experience in creating production-quality ML code.
• Familiarity with modern ML frameworks (e.g., PyTorch, JAX, TensorFlow) and model libraries/tooling (e.g., Hugging Face ecosystem, vLLM, distributed training stacks).
• Experience in designing and implementing evaluation pipelines for LLM/ML systems, including metrics computation, dataset management, and experiment comparisons.
• Strong understanding of data pipelines and ML systems engineering, emphasizing reproducibility, observability, and debugging.
• Experience with large-scale distributed ML systems and performance optimization for training and evaluation workloads (preferably in GPU/accelerator environments).
• Experience with large-scale data processing and workflow orchestration to support model training and evaluation.
• Ability to collaborate directly with technical stakeholders, including research scientists, ML engineers, data engineers, and customer technical leads.
• Excellent written and verbal communication skills, including the capacity to explain complex technical trade-offs to both technical and non-technical audiences.
• Competitive salary and performance-based incentives.
• Opportunities for professional development and continuous learning.
• Flexible work arrangements and a supportive work environment.
• Health, dental, and vision insurance options.
• Generous paid time off and holiday policies.
Guardian Industries - DeWitt
EVERSANA
Nagarro
Get handpicked remote jobs straight to your inbox weekly.