
AI Response Labeler – Italian Specialty
Posted Sep 9

Posted Sep 9
This is a fully remote position, open to applicants in Peru.
• Conduct side-by-side evaluations of AI-generated responses to identify the stronger option.
• Assess responses based on criteria such as factual accuracy, relevance, completeness, clarity, reasoning, adherence to instructions, tone, and overall quality.
• Evaluate content provided in English, Italian, or a combination of both languages.
• Analyze general-purpose inquiries and answers, web search results, file-based tasks, image-based responses, content generation requests, and both single-turn and multi-turn dialogues.
• Utilize expertise in Italian language, terminology, tone, regional nuances, idioms, and cultural context unique to Italy.
• Focus on the overall quality of responses rather than solely on grammar, translation, or fluency in language.
• Identify unsupported assertions, incomplete reasoning, overlooked instructions, unnatural phrasing, cultural inaccuracies, and variances in usefulness.
• Accurately and consistently apply detailed, scenario-specific annotation guidelines.
• Make independent evaluative decisions in cases of ambiguity.
• Document decisions and provide clear, evidence-based rationale when necessary.
• Complete evaluations within specified time and productivity targets.
• Maintain consistent judgment across a high volume of diverse assignments.
• Engage in training, guided practice, calibration sessions, qualification reviews, and ongoing quality assessment activities.
• Integrate feedback and adjust evaluation judgments to align with team and client quality standards.
• Achieve a minimum of 25 evaluation tasks daily, with most tasks estimated to take around 15 minutes each.
• Successfully finish the structured onboarding and qualification program prior to starting production work.
• Native-level or professional proficiency in Italian.
• Extensive familiarity with Italian linguistic conventions, regional vocabulary, idioms, tone, and cultural context as they are used in Italy.
• Strong command of English and excellent reading comprehension skills.
• Robust analytical and critical-thinking abilities.
• Capability to evaluate content across a variety of subjects, formats, and task types.
• Proficiency in assessing factual accuracy, relevance, reasoning, clarity, adherence to instructions, cultural appropriateness, and overall usefulness.
• Skill in recognizing nuanced differences in meaning, quality, tone, and user intent.
• Sound judgment in applying structured evaluation criteria to ambiguous or unfamiliar situations.
• Strong written communication skills, enabling clear and concise explanations of evaluation decisions.
• Exceptional attention to detail with the ability to maintain accuracy within set time frames.
• Ability to learn and consistently utilize detailed evaluation frameworks.
• Capacity to work independently while remaining aligned with collective quality standards.
• Comfort in performing repetitive, detail-oriented tasks for extended durations while preserving focus, accuracy, and consistent judgment.
• Openness to receiving feedback, recalibrating decisions, and adapting as evaluation guidelines evolve.
• Preferred experience in side-by-side labeling, annotation, comparative content evaluation, quality assessment, AI-generated response evaluation, data labeling, search relevance, content quality, factual accuracy, user-facing digital experiences, or detailed guidelines/rubrics.
• Medical, dental, and vision insurance.
• Flexible Spending Account (FSA).
• 401(k) retirement plan.
• Competitive paid time off.
• Parental leave.
• Opportunities for professional growth and development.
• Benefits in accordance with local regulations and the terms of employment through an Employer of Record partner.
Mercor
BPCS, Comprehensive marketing solutions, ltd.
Tech Mahindra
Get handpicked remote jobs straight to your inbox weekly.