AI Quality & Evaluation Lead

Posted 1 day ago

This is a fully remote position, open to applicants in Poland.

📋 Description

• Develop the quality evaluation framework for AI-driven coaching within Samsung Health.

• Create a taxonomy of failures for coaching outputs.

• Manually assess a minimum of 150 traces with open-coded annotations, which should include at least 40 synthetic profiles with sparse data.

• Classify failures into 5 to 10 distinct categories and measure their occurrence.

• Set aside at least 50 traces as a holdout sample.

• Establish binary pass/fail criteria for each category in the taxonomy.

• Draft annotation guidelines that include both pass and fail examples.

• Double-code at least 30 traces, resolve any discrepancies, and keep a log of guideline revisions.

• Generate judge prompts for every evaluation criterion.

• Assess the LLM judge by calculating true-positive and true-negative rates on both development and holdout datasets.

• Conduct weekly updates on aspects like helpfulness, relevance, and tone.

• Document a revalidation process that is triggered by any model or prompt changes, as well as quarterly assessments.

• Create a comprehensive playbook detailing grading processes, taxonomy updates, rubric modifications, judge revalidation, and weekly reviews.

• Perform a handover test to enable the internal owner to independently rerun validation and weekly reviews.

• Secure approval from the Head of Product on the taxonomy and rubric.

• Work collaboratively with an internal owner to transfer the methodology.


⛳️ Requirements

• Proven experience in executing the complete evaluation loop at least once on conversational or generated-text outputs.

• Familiarity with error analysis on actual traces.

• Experience in developing a failure taxonomy.

• Ability to create binary criteria and annotation guidelines.

• Experience in validating an LLM judge against personal labels.

• Background in conversation design, AI product quality, model behavior/policy, content design, human data operations, UX research with a strong focus on qualitative coding, RLHF, applied linguistics, or product management.

• Skill in identifying previously unrecognized failure modes through output analysis.

• Capacity to serve as an arbiter, documenting overrulings and the rationale behind them.

• Recognition that rubrics are developed through the grading process.

• Ability to communicate the limitations of judge agreement rates.

• Proficiency in writing clear and unambiguous annotation guidelines.

• Comfort with working in notebooks and spreadsheets.

• No production coding experience is necessary.

• Expertise in nutrition, weight management, or behavior change is not required.

• Must be willing to manage personal taxes and understand that there is no paid time off.


🏝️ Benefits

• Remote-first work environment.

• Independent contractor arrangement.

• Flexible remote working options.

• A global team of over 100 individuals across more than 30 countries.

People also viewed

The Hershey Company16 hours ago

Senior Manager, AI Business Modernization

US flagUnited States OnlyFull-timeArtificial Intelligence$143.1k – $178.9k/year
ApplyView job
TaskUs16 hours ago

Head of Delivery, AI Operations

US flagUnited States OnlyFull-timeArtificial Intelligence$156k – $234k/year
ApplyView job
Supabase16 hours ago

Head of AI Native Operations

Anywhere in the WorldFull-timeArtificial Intelligence
ApplyView job
321 The Agency17 hours ago

Director of Technology, AI

US flagFlorida OnlyFull-timeArtificial Intelligence$100k/year
ApplyView job
neuefische GmbH - School and Pool for Digital Talent20 hours ago

Freelance Coach – AI-Powered Marketing, German-Speaking

DE flagGermany OnlyFreelanceArtificial Intelligence
ApplyView job
Williams Lea1 day ago

AI Workflow Engineer

GB flagUnited Kingdom OnlyFull-timeArtificial Intelligence£60k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers