
AI Safety Expert – English, Danish
Posted Sep 15

Posted Sep 15
This is a fully remote position, open to applicants in United States.
• Conduct red team assessments on conversational AI models and agents through methods such as jailbreaks, prompt injections, misuse scenarios, bias exploitation, and multi-turn manipulation.
• Create human data by annotating errors, categorizing vulnerabilities, and identifying systemic risks.
• Implement taxonomies, benchmarks, and playbooks to ensure consistent testing practices.
• Generate reproducible reports, datasets, and case studies of attacks.
• Assess AI outputs that pertain to sensitive subjects like bias, misinformation, or harmful behaviors.
• Identify vulnerabilities that automated testing may overlook.
• Enhance evaluation coverage and minimize unexpected issues during production.
• Participate in projects aimed at training and improving AI systems for Mercor's AI lab and its enterprise clients.
• Proficiency in both English and Danish is mandatory.
• Previous experience in red teaming within AI adversarial contexts, cybersecurity, or socio-technical exploration is essential.
• Capability to adversarially probe AI systems and challenge them to their limits.
• Familiarity with frameworks or benchmarks for organized testing.
• Skill in clearly communicating risks to both technical and non-technical audiences.
• Flexibility to adapt across various projects and client needs.
• Engagement as an independent contractor.
• Must not require H1-B or STEM OPT visa sponsorship.
• Preferred: experience in adversarial machine learning, including knowledge of jailbreak datasets, prompt injection, RLHF/DPO attacks, or model extraction.
• Preferred: background in cybersecurity, encompassing penetration testing, exploit development, or reverse engineering.
• Preferred: experience in socio-technical risk, including harassment/disinformation probing, abuse analysis, or conversational AI testing.
• Preferred: creative probing skills involving psychology, acting, or writing.
• Fully remote position.
• Flexible work schedule that you set yourself.
• Weekly payments through Stripe or Wise.
• Opportunity to gain experience in human data-driven AI red teaming.
• Play a direct role in enhancing the robustness, safety, and trustworthiness of AI systems.
• Competitive compensation.
• Collaboration with top researchers in the field.
• Access to wellness resources and clear protocols for higher-sensitivity projects.
• Reasonable accommodations available upon request.
• Referral bonuses of up to $250 for each successful referral.
The College Board
Mercor
Mercor
Get handpicked remote jobs straight to your inbox weekly.