Senior Consultant, AI Safety

Posted Aug 22

This is a fully remote position, open to applicants in United Kingdom, +3 more countries.

📋 Description

• Perform red teaming and adversarial assessments of AI systems against established harm categories.

• Evaluate model responses based on harm and risk criteria, providing expert analysis.

• Utilize subject matter expertise in areas such as grooming and CSEA, radicalization pathways, crisis signaling, or teen online safety.

• Assist in the creation of evaluation frameworks that convert real-world harm knowledge into structured, testable criteria.

• Contribute to intervention logic that connects at-risk users with appropriate support.

• Prepare methodologies or findings tailored for technical and government audiences.

• Collaborate closely with Moonshot's AI Safety team to advance its AI Safety portfolio.

• Ensure compliance with contractual, legal, data protection, and ethical obligations in all deliverables.

• This role does not include responsibilities related to client relationship management, team leadership, or business development.


⛳️ Requirements

• Proven experience in trust and safety, online harms, or a closely related area such as violence prevention, safeguarding, or public health, with the ability to apply this expertise to AI systems.

• Documented experience in designing research, evaluation frameworks, or interventions for harm categories including violent extremism, CSEA, self-harm and crisis, or targeted violence.

• Capability to convert real-world knowledge of harm mechanisms into methods for testing AI system safety.

• Comfort and resilience when handling highly sensitive or graphic content, including violence, extremist material, and crisis situations, with an understanding of well-being practices for this type of work.

• Excellent written communication skills, capable of producing credible, non-promotional content for technical and government audiences.

• Strong judgment in navigating ambiguity and sensitive materials.

• Availability for nearly full-time engagement over an estimated 6-week period.

• Willingness to undergo necessary security clearance procedures, if required by the engagement.

• Desirable: substantial knowledge in model safety, red teaming, or adversarial evaluation of LLMs or other AI systems.

• Desirable: familiarity with LLM architecture, safety tools, or trust and safety policies.

• Desirable: involvement in child safety evaluation, teen safety product development, or grooming and CSEA detection.

• Desirable: design experience in intervention or diversion programs that can be adapted to AI-mediated interventions.

• Desirable: engagement with government or regulatory bodies, including briefing officials or supporting policy submissions.

• Desirable: academic or practical background in radicalization studies, forensic psychology, or violence risk assessment.

• Desirable: experience in developing taxonomies or classifiers, including how testing data informs classifier development.


🏝️ Benefits

• Flexible working arrangements.

• Opportunity to engage in diverse and impactful projects.

• Competitive consultancy rates.

• Options for remote working available.

People also viewed

Mercor1 day ago

AI Safety Red Teamer

US flagUnited States OnlyFreelanceArtificial Intelligence$70 – $84/hour
ApplyView job
Mercor1 day ago

AI Safety Red Teamer

US flagUnited States OnlyFreelanceArtificial Intelligence$70 – $84/hour
ApplyView job
Mercor1 day ago

AI Safety Expert, English – Kannada

US flagUnited States OnlyFreelanceArtificial Intelligence$16 – $22/hour
ApplyView job
Mercor1 day ago

AI Safety Expert – English, Punjabi

US flagUnited States OnlyFreelanceArtificial Intelligence$16 – $22/hour
ApplyView job
Mercor1 day ago

AI Safety Expert – English, Tamil

US flagUnited States OnlyFreelanceArtificial Intelligence$16 – $22/hour
ApplyView job
Mercor1 day ago

AI Safety Expert – English, Tamil

US flagUnited States OnlyFreelanceArtificial Intelligence$16 – $22/hour
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers