
Senior Consultant, AI Safety
Posted Aug 22

Posted Aug 22
This is a fully remote position, open to applicants in United Kingdom, +3 more countries.
• Perform red teaming and adversarial assessments of AI systems against established harm categories.
• Evaluate model responses based on harm and risk criteria, providing expert analysis.
• Utilize subject matter expertise in areas such as grooming and CSEA, radicalization pathways, crisis signaling, or teen online safety.
• Assist in the creation of evaluation frameworks that convert real-world harm knowledge into structured, testable criteria.
• Contribute to intervention logic that connects at-risk users with appropriate support.
• Prepare methodologies or findings tailored for technical and government audiences.
• Collaborate closely with Moonshot's AI Safety team to advance its AI Safety portfolio.
• Ensure compliance with contractual, legal, data protection, and ethical obligations in all deliverables.
• This role does not include responsibilities related to client relationship management, team leadership, or business development.
• Proven experience in trust and safety, online harms, or a closely related area such as violence prevention, safeguarding, or public health, with the ability to apply this expertise to AI systems.
• Documented experience in designing research, evaluation frameworks, or interventions for harm categories including violent extremism, CSEA, self-harm and crisis, or targeted violence.
• Capability to convert real-world knowledge of harm mechanisms into methods for testing AI system safety.
• Comfort and resilience when handling highly sensitive or graphic content, including violence, extremist material, and crisis situations, with an understanding of well-being practices for this type of work.
• Excellent written communication skills, capable of producing credible, non-promotional content for technical and government audiences.
• Strong judgment in navigating ambiguity and sensitive materials.
• Availability for nearly full-time engagement over an estimated 6-week period.
• Willingness to undergo necessary security clearance procedures, if required by the engagement.
• Desirable: substantial knowledge in model safety, red teaming, or adversarial evaluation of LLMs or other AI systems.
• Desirable: familiarity with LLM architecture, safety tools, or trust and safety policies.
• Desirable: involvement in child safety evaluation, teen safety product development, or grooming and CSEA detection.
• Desirable: design experience in intervention or diversion programs that can be adapted to AI-mediated interventions.
• Desirable: engagement with government or regulatory bodies, including briefing officials or supporting policy submissions.
• Desirable: academic or practical background in radicalization studies, forensic psychology, or violence risk assessment.
• Desirable: experience in developing taxonomies or classifiers, including how testing data informs classifier development.
• Flexible working arrangements.
• Opportunity to engage in diverse and impactful projects.
• Competitive consultancy rates.
• Options for remote working available.
Mercor
Mercor
Mercor
Mercor
Get handpicked remote jobs straight to your inbox weekly.