
AI Quality & Evaluation Lead
Posted 1 day ago

Posted 1 day ago
This is a fully remote position, open to applicants in Poland.
• Develop the quality evaluation framework for AI-driven coaching within Samsung Health.
• Create a taxonomy of failures for coaching outputs.
• Manually assess a minimum of 150 traces with open-coded annotations, which should include at least 40 synthetic profiles with sparse data.
• Classify failures into 5 to 10 distinct categories and measure their occurrence.
• Set aside at least 50 traces as a holdout sample.
• Establish binary pass/fail criteria for each category in the taxonomy.
• Draft annotation guidelines that include both pass and fail examples.
• Double-code at least 30 traces, resolve any discrepancies, and keep a log of guideline revisions.
• Generate judge prompts for every evaluation criterion.
• Assess the LLM judge by calculating true-positive and true-negative rates on both development and holdout datasets.
• Conduct weekly updates on aspects like helpfulness, relevance, and tone.
• Document a revalidation process that is triggered by any model or prompt changes, as well as quarterly assessments.
• Create a comprehensive playbook detailing grading processes, taxonomy updates, rubric modifications, judge revalidation, and weekly reviews.
• Perform a handover test to enable the internal owner to independently rerun validation and weekly reviews.
• Secure approval from the Head of Product on the taxonomy and rubric.
• Work collaboratively with an internal owner to transfer the methodology.
• Proven experience in executing the complete evaluation loop at least once on conversational or generated-text outputs.
• Familiarity with error analysis on actual traces.
• Experience in developing a failure taxonomy.
• Ability to create binary criteria and annotation guidelines.
• Experience in validating an LLM judge against personal labels.
• Background in conversation design, AI product quality, model behavior/policy, content design, human data operations, UX research with a strong focus on qualitative coding, RLHF, applied linguistics, or product management.
• Skill in identifying previously unrecognized failure modes through output analysis.
• Capacity to serve as an arbiter, documenting overrulings and the rationale behind them.
• Recognition that rubrics are developed through the grading process.
• Ability to communicate the limitations of judge agreement rates.
• Proficiency in writing clear and unambiguous annotation guidelines.
• Comfort with working in notebooks and spreadsheets.
• No production coding experience is necessary.
• Expertise in nutrition, weight management, or behavior change is not required.
• Must be willing to manage personal taxes and understand that there is no paid time off.
• Remote-first work environment.
• Independent contractor arrangement.
• Flexible remote working options.
• A global team of over 100 individuals across more than 30 countries.
The Hershey Company
TaskUs
Supabase
321 The Agency
Get handpicked remote jobs straight to your inbox weekly.