
AI Safety Expert, English, Dutch
Posted Sep 3

Posted Sep 3
This is a fully remote position, open to applicants in United States.
β’ Conduct red teaming on conversational AI models and agents by exploring jailbreaks, prompt injections, misuse scenarios, bias exploitation, and multi-turn manipulation.
β’ Create human data through annotating failures, categorizing vulnerabilities, and identifying systemic risks.
β’ Utilize taxonomies, benchmarks, and playbooks to ensure consistent testing procedures.
β’ Generate reproducible reports, datasets, and attack scenarios for clients.
β’ Analyze AI outputs related to sensitive issues like bias, misinformation, or harmful behaviors.
β’ Identify vulnerabilities that automated testing may overlook.
β’ Broaden evaluation coverage and minimize unexpected outcomes in production.
β’ Assist Mercor clients in enhancing the safety, robustness, and trustworthiness of their AI systems.
β’ Native fluency in both English and Dutch is essential.
β’ Previous experience in red teaming within AI adversarial contexts, cybersecurity, or socio-technical probing is required.
β’ Capability to probe systems adversarially and push them to their limits.
β’ Experience with frameworks or benchmarks for structured testing is necessary.
β’ Proficient in articulating risks to both technical and non-technical stakeholders.
β’ Flexibility to adapt to various projects and clients.
β’ Nice-to-have: Experience in adversarial machine learning, including jailbreak datasets, prompt injection, RLHF/DPO attacks, or model extraction.
β’ Nice-to-have: Experience in cybersecurity, such as penetration testing, exploit development, or reverse engineering.
β’ Nice-to-have: Experience with socio-technical risks, including harassment/disinformation probing, abuse analysis, or testing conversational AI.
β’ Nice-to-have: Creative probing experience in fields like psychology, acting, or writing.
β’ Must be engaged as an independent contractor.
β’ H1-B and STEM OPT candidates are not supported.
β’ Fully remote position.
β’ Flexible work schedule.
β’ Weekly payments via Stripe or Wise based on delivered services.
β’ Project timelines can be adjusted, whether extended, shortened, or concluded early based on requirements and performance.
β’ Participation in higher-sensitivity projects is optional.
β’ Clear guidelines and wellness resources available for handling sensitive content.
β’ Competitive compensation.
β’ Referral bonus of up to $250 for each successful referral.
The College Board
Mercor
Mercor
Get handpicked remote jobs straight to your inbox weekly.