
AI Response Labeler – Annotator
Posted 7 hours ago

Posted 7 hours ago
This is a fully remote position, open to applicants in Latin America.
• Conduct side-by-side evaluations of AI-generated responses to determine the superior option based on established assessment criteria.
• Evaluate responses for accuracy, relevance, completeness, reasoning, adherence to instructions, clarity, safety, tone, and overall usefulness.
• Analyze questions and answers, web-search results, file-based and image-based responses, content creation, as well as single-turn or multi-turn conversations.
• Detect unsupported claims, overlooked instructions, weak reasoning, and incomplete responses.
• Utilize detailed, scenario-specific guidelines and make decisions when examples do not yield a clear answer.
• Provide concise, evidence-based justifications for evaluation choices when necessary.
• Achieve productivity targets while ensuring accuracy and consistent judgment.
• Engage in training, guided practice, calibration, qualification reviews, and ongoing quality assessments.
• Adapt to feedback as evaluation guidelines and quality standards progress.
• Generally complete at least 25 evaluation tasks daily, with the average task duration being around 15 minutes.
• Exceptional written English comprehension and communication capabilities, including the ability to interpret complex prompts and guidelines and articulate evaluation decisions clearly.
• Strong critical-thinking abilities and the capacity to evaluate content across a diverse array of subjects.
• Sound judgment when assessing factual accuracy, reasoning, user intent, and nuanced variations in response quality.
• Meticulous attention to detail and the ability to consistently apply structured criteria.
• Comfort with repetitive, focused tasks and managing a high volume of evaluations.
• Capability to work independently, accept feedback, and remain aligned with shared quality standards.
• During the approximately 30-day training and qualification phase, employees are required to work from 9:00 a.m. to 5:00 p.m. Pacific Time.
• All new hires must successfully complete a structured onboarding and qualification program prior to commencing production work.
• Preferred: experience in evaluating, ranking, or comparing AI-generated responses.
• Preferred: background in data annotation, content quality assessment, search relevance evaluation, or model-quality review.
• Preferred: familiarity with detailed rubrics, annotation guidelines, or quality benchmarks.
• Preferred: experience in writing clear rationales that support evaluation decisions.
• Medical, dental, and vision coverage
• Flexible Spending Account (FSA)
• 401(k) retirement plan
• Competitive paid time off
• Parental leave
• Professional growth and development opportunities
Mercor
Tech Mahindra
PwC
Get handpicked remote jobs straight to your inbox weekly.