
Machine Learning Engineer – Model Evaluation, Experimentation
Posted Aug 10

Posted Aug 10
This is a fully remote position, open to applicants in United States.
• Create well-structured, multi-phase machine learning tasks inspired by genuine ML research concepts.
• Implement modifications, conduct training experiments, and assess outcomes to determine effective solutions.
• Develop tasks centered around reinforcement learning principles, including reward functions and training behaviors.
• Analyze cutting-edge models and pinpoint areas where they underperform and the reasons behind it.
• Collaborate with researchers and peers to ensure consistency, rigor, and impartiality in the work.
• Engage in a close feedback loop with the lab’s researchers.
• MSc or PhD in machine learning, computer science, or another STEM discipline, or equivalent practical experience in a research-intensive area.
• Over 1 year of experience in a research or research-engineering capacity.
• Practical experience in training and assessing ML models and conducting experiments from setup through execution to analysis.
• Strong knowledge of large language models, including their strengths, weaknesses, and evaluation methods.
• Proficient in Python and Git.
• Comfortable working in scripting and notebook environments.
• A basic grasp of reinforcement learning concepts, including reward functions and policy training, is preferred.
• Previous experience in AI training, model assessment, or benchmark/task authorship is preferred.
• High level of attention to detail and creativity in designing tasks.
• Excellent written communication skills.
• Ability to independently navigate ambiguous and open-ended challenges.
• Availability for approximately 35 hours per week.
• W-2 employment.
• Payroll, benefits, and compliance managed by Cincinnatus LLC.
• Fully remote work opportunity within the United States.
• Approximately 35 hours of work per week.
• Chance to be placed at a top-tier AI lab as part of its extended workforce.
• Reasonable accommodations available for qualified individuals with disabilities and disabled veterans throughout the application process.
Thrive Market
BlueFlag LLP
Shield AI
Weekday (YC W21)
Get handpicked remote jobs straight to your inbox weekly.