
AI Safety Expert, English β Dutch
Posted Aug 10

Posted Aug 10
This is a fully remote position, open to applicants in United States.
β’ Conduct red team assessments on conversational AI models and agents utilizing jailbreaks, prompt injections, misuse scenarios, bias exploitation, and multi-turn manipulations.
β’ Create human data by annotating failures, categorizing vulnerabilities, and identifying systemic risks.
β’ Implement taxonomies, benchmarks, and playbooks to maintain consistency in testing.
β’ Generate reproducible reports, datasets, and attack scenarios for clients.
β’ Examine AI outputs related to sensitive topics, including bias, misinformation, and harmful behaviors.
β’ Reveal vulnerabilities that automated tests may overlook.
β’ Broaden evaluation coverage across additional scenarios.
β’ Enhance customer AI systems to improve their safety, robustness, and trustworthiness.
β’ Collaborate with leading researchers on projects aimed at training and refining AI systems.
β’ Previous red teaming experience in AI adversarial work, cybersecurity, or socio-technical probing.
β’ Native proficiency in English and Dutch.
β’ Experience with adversarial inputs and evaluating AI models for vulnerabilities.
β’ Ability to utilize frameworks, taxonomies, benchmarks, or playbooks for structured testing procedures.
β’ Capability to articulate risks clearly to both technical and non-technical stakeholders.
β’ Flexibility to adapt across various projects and clients.
β’ H1-B and STEM OPT candidates are not eligible.
β’ Nice-to-have: experience in adversarial machine learning, including jailbreak datasets, prompt injection, RLHF/DPO attacks, or model extraction.
β’ Nice-to-have: cybersecurity experience, such as penetration testing, exploit development, or reverse engineering.
β’ Nice-to-have: socio-technical risk experience, including harassment/disinformation probing, abuse analysis, or conversational AI testing.
β’ Nice-to-have: creative probing experience in psychology, acting, or writing.
β’ Fully remote work.
β’ Flexible schedule with the option to work on your own terms.
β’ Weekly payments through Stripe or Wise.
β’ Participation in higher-sensitivity projects is optional.
β’ Clear guidelines and wellness resources available for higher-sensitivity projects.
β’ Competitive compensation.
β’ Opportunity to collaborate with leading researchers.
β’ Referral bonus of up to $250 for each successful referral.
The College Board
Mercor
Mercor
Get handpicked remote jobs straight to your inbox weekly.