
AI Policy Generalist
Posted 14 hours ago

Posted 14 hours ago
This is a fully remote position, open to applicants in United States.
• Acquire knowledge of customer policies, definitions, taxonomies, and evaluation rubrics.
• Assess user inquiries and AI model outputs within the complete context of conversations.
• Categorize cases according to the most relevant policy category.
• Differentiate between closely related labels, severity levels, and policy boundaries.
• Choose defensible classifications for ambiguous situations.
• Compose clear, evidence-based rationales that reference policy language and conversation specifics.
• Detect gaps in policy, contradictions, vague definitions, and emerging edge cases.
• Raise inquiries when existing guidance fails to address a case.
• Engage in calibration discussions with evaluators, project leaders, policy teams, and researchers.
• Consistently apply customer policies without letting personal beliefs interfere.
• Ensure precision and attention to detail across repeated evaluations.
• Integrate feedback and implement clarified guidance.
• Contribute to the enhancement of evaluation frameworks, examples, decision rules, and quality standards.
• Transition between projects that address various policy domains and customer requirements.
• Must reside in the United States.
• Must possess lawful authorization to work in the United States for Handshake.
• Must be available to work from Monday to Friday between 8AM and 5PM PT as core hours.
• Ability to engage thoughtfully, responsibly, and sustainably with sensitive material, including sexual content, emotional distress, self-harm, suicide, violence, weapons, abuse, exploitation, and discrimination.
• Strong judgment and consistent quality of work.
• Capability to rapidly learn unfamiliar subjects.
• Proficiency in clear and precise written communication.
• Ability to apply detailed policies, definitions, taxonomies, rubrics, and evaluation frameworks.
• Skill in distinguishing closely related labels, severity levels, and policy boundaries.
• Capability to write concise, evidence-based rationales.
• Ability to participate in calibration discussions and assimilate feedback.
• Previous experience in AI evaluation is beneficial but not mandatory.
• Strong candidates may have backgrounds in quality assurance, research, editing, law, teaching, operations, trust and safety, content moderation, social science, policy, investigations, compliance, or customer support.
• W-2 employment classification.
• Monday to Friday work schedule.
• Core hours from 8AM to 5PM PT.
• Remote work arrangement.
10a Labs
Globalme (now Summa Linguae Technologies)
UserTesting
Mentor Collective
Get handpicked remote jobs straight to your inbox weekly.