
AI Response Labeler – Italian Specialty
Posted Sep 9

Posted Sep 9
This is a fully remote position, open to applicants in Brazil.
• Conduct side-by-side evaluations of AI-generated replies to identify the stronger response.
• Assess responses for accuracy, relevance, completeness, clarity, reasoning, adherence to instructions, tone, and overall quality.
• Review content in English, Italian, or a mix of both languages.
• Evaluate general inquiries and answers, web search results, file-based tasks, image responses, content generation requests, and both single-turn and multi-turn dialogues.
• Utilize expertise in Italian language, terminology, tone, regional norms, idioms, and cultural context specific to Italy.
• Focus on the quality of complete responses rather than only grammar, translation, or language fluency.
• Detect unsupported assertions, incomplete reasoning, overlooked instructions, awkward phrasing, cultural inaccuracies, and variations in usefulness.
• Accurately and consistently apply detailed, scenario-specific annotation guidelines.
• Make independent evaluation decisions when the guidelines do not provide clear answers.
• Record decisions and offer concise, evidence-based justifications when needed.
• Fulfill evaluations within set time and productivity benchmarks.
• Maintain consistent judgment across a large volume of diverse assignments.
• Engage in training, guided practice, calibration sessions, qualification reviews, and ongoing quality assurance activities.
• Integrate feedback and adjust evaluation conclusions to meet team and client quality standards.
• Complete at least 25 evaluation tasks daily, with most tasks taking around 15 minutes.
• Native-level or professional proficiency in Italian.
• Extensive knowledge of Italian linguistic norms, regional vocabulary, idioms, tone, and cultural context pertaining to Italy.
• Strong command of English and excellent reading comprehension skills.
• Ability to comprehend complex prompts, AI-generated responses, and detailed annotation guidelines written in English.
• Strong analytical and critical thinking capabilities.
• Ability to evaluate content across a range of topics, formats, and task types.
• Competence in assessing factual accuracy, relevance, reasoning, clarity, adherence to instructions, cultural appropriateness, and overall usefulness.
• Capacity to identify subtle differences in meaning, quality, tone, and user intent.
• Sound judgment in applying structured evaluation criteria to ambiguous or unfamiliar situations.
• Excellent written communication skills.
• High attention to detail and the ability to maintain accuracy within set time constraints.
• Ability to learn and consistently implement detailed evaluation frameworks.
• Capability to work independently while adhering to shared quality standards.
• Comfort with performing repetitive, detail-oriented tasks for extended periods while sustaining focus, accuracy, and consistent judgment.
• Ability to accept feedback, recalibrate decisions, and adapt as evaluation guidelines evolve.
• Preferred experience in side-by-side labeling, annotation, comparative content evaluation, quality assessment, AI-generated response evaluation, model quality assessment, data labeling or annotation, search relevance, content quality, factual accuracy, user-facing digital experiences, or structured decision-making frameworks.
• Must successfully complete structured onboarding and qualification before starting production work.
• Medical, dental, and vision insurance.
• Flexible Spending Account (FSA).
• 401(k) retirement plan.
• Competitive paid time off.
• Parental leave.
• Opportunities for professional growth and development.
• Benefits aligned with local regulations and terms of employment through an Employer of Record partner.
Mercor
BPCS, Comprehensive marketing solutions, ltd.
Tech Mahindra
Get handpicked remote jobs straight to your inbox weekly.