
AI Evaluation Specialist
Posted Sep 13

Posted Sep 13
This is a fully remote position, open to applicants in United States, +6 more locations.
• Evaluate outputs generated by AI against comprehensive rubrics, guidelines, and established quality standards.
• Assess responses for precision, relevance, completeness, reasoning quality, and compliance with instructions.
• Apply consistent and unbiased judgment across extensive volumes of evaluation examples.
• Identify outputs that do not meet essential quality or task requirements.
• Maintain dependable assessment standards throughout repeated evaluation processes.
• Spot reasoning gaps, logical errors, inconsistencies, unsupported conclusions, and failures in tool use.
• Analyze deviations of AI-generated outputs from anticipated reasoning or quality benchmarks.
• Document recurring weaknesses in models and opportunities for enhancement.
• Generate clear, concise, and actionable written feedback on strengths and areas needing improvement.
• Clarify the rationale behind evaluation decisions and quality ratings.
• Ensure detailed, transparent, traceable, and reproducible assessment documentation.
• Engage in discussions regarding rubric interpretation and ambiguous cases.
• Assist in refining assessment criteria as AI models and project specifications evolve.
• Provide insights to support process optimization and best practices in evaluation.
• Collaborate with fellow reviewers to enhance alignment and reliability across evaluation workflows.
• Proven experience in grading, quality assurance, editorial review, assessment, annotation, or any other field necessitating meticulous analysis and detailed feedback.
• Advanced and regular use of AI assistants such as ChatGPT, Claude, or equivalent tools for professional tasks and productivity.
• Strong capability to synthesize intricate information and clearly communicate conclusions in writing.
• Experience in process improvement, rubric development, operational quality assessment, or structured evaluation workflows is a plus.
• Strong critical thinking skills with a particular focus on consistency, integrity, and fairness.
• High attention to detail and comfort in reviewing large sets of similar examples.
• Ability to work independently while ensuring consistent evaluation quality.
• Collaborative mindset for discussing ambiguous cases and refining common assessment standards.
• Excellent written English and professional documentation capabilities.
• Based in the United States, Canada, United Kingdom, Ireland, Australia, or New Zealand.
• Authorized to perform contract work in the relevant country.
• No previous formal experience in AI research or model training is necessary.
• Engagement as a part-time independent contractor.
• Fully remote work environment.
• Flexible project scope, workload, timing, and duration based on project needs.
Mercor
Mercor
Mercor
Mercor
Get handpicked remote jobs straight to your inbox weekly.