
Machine Learning Engineer – Model Evaluation, Experimentation
Posted Jul 30

Posted Jul 30
This is a fully remote position, open to applicants in United States.
• Create realistic benchmark tasks for machine learning that are based on research workflows, encompassing model implementation, experimentation, training, evaluation, and performance analysis.
• Convert open-ended research ideas into structured, reproducible evaluation tasks with well-defined success criteria.
• Develop machine learning solutions utilizing Python, conduct experiments, and generate reference implementations that illustrate correct methodologies and anticipated outcomes.
• Create benchmark tasks that incorporate reinforcement learning concepts such as reward functions, policy optimization, training dynamics, and model behavior when relevant.
• Assess AI-generated solutions by detecting implementation errors, experimental flaws, incorrect reasoning, and unsupported conclusions.
• Work collaboratively with AI researchers and other subject matter experts to enhance benchmark quality, technical rigor, and evaluation consistency.
• A Master's degree, PhD, or equivalent practical experience in Machine Learning, Computer Science, Artificial Intelligence, Data Science, or another quantitative STEM field.
• At least 1 year of professional experience in machine learning research, research engineering, applied AI, or a similar research-intensive technical position.
• Strong hands-on experience in designing, training, evaluating, and optimizing machine learning models through comprehensive experimental workflows.
• Practical experience in conducting machine learning experiments, including experiment setup, hyperparameter tuning, execution, validation, and analysis.
• A solid understanding of modern Large Language Models (LLMs), including their capabilities, limitations, and evaluation methodologies.
• Proficiency in Python and Git, with experience in both script-based and notebook-based development environments.
• Familiarity with reinforcement learning concepts—such as reward functions, policy optimization, and training behavior—is preferred.
• Experience in AI evaluation, benchmark development, AI training, or task authoring is highly desirable.
• Exceptional analytical thinking, creativity, attention to detail, and the ability to independently tackle complex, open-ended technical problems.
• Strong written communication skills for documenting experimental methodologies and technical findings.
• Ability to consistently commit approximately 35 hours per week.
• Competitive salary and comprehensive benefits package.
• Opportunities for professional development and continuous learning.
• Collaborative and innovative work environment.
• Flexible work schedule and remote work options.
Doma
CSC Generation
Accelerant
Capgemini
Get handpicked remote jobs straight to your inbox weekly.