
AI Safety Expert, English – Dutch
Posted 1 day ago

Posted 1 day ago
This is a fully remote position, open to applicants in United States.
• Conduct red team assessments on conversational AI models and agents utilizing jailbreaks, prompt injections, misuse scenarios, bias exploitation, and multi-turn manipulations.
• Create human data by annotating failures, categorizing vulnerabilities, and identifying systemic risks.
• Implement taxonomies, benchmarks, and playbooks to maintain consistency in testing.
• Generate reproducible reports, datasets, and attack scenarios for clients.
• Examine AI outputs related to sensitive topics, including bias, misinformation, and harmful behaviors.
• Reveal vulnerabilities that automated tests may overlook.
• Broaden evaluation coverage across additional scenarios.
• Enhance customer AI systems to improve their safety, robustness, and trustworthiness.
• Collaborate with leading researchers on projects aimed at training and refining AI systems.
• Previous red teaming experience in AI adversarial work, cybersecurity, or socio-technical probing.
• Native proficiency in English and Dutch.
• Experience with adversarial inputs and evaluating AI models for vulnerabilities.
• Ability to utilize frameworks, taxonomies, benchmarks, or playbooks for structured testing procedures.
• Capability to articulate risks clearly to both technical and non-technical stakeholders.
• Flexibility to adapt across various projects and clients.
• H1-B and STEM OPT candidates are not eligible.
• Nice-to-have: experience in adversarial machine learning, including jailbreak datasets, prompt injection, RLHF/DPO attacks, or model extraction.
• Nice-to-have: cybersecurity experience, such as penetration testing, exploit development, or reverse engineering.
• Nice-to-have: socio-technical risk experience, including harassment/disinformation probing, abuse analysis, or conversational AI testing.
• Nice-to-have: creative probing experience in psychology, acting, or writing.
• Fully remote work.
• Flexible schedule with the option to work on your own terms.
• Weekly payments through Stripe or Wise.
• Participation in higher-sensitivity projects is optional.
• Clear guidelines and wellness resources available for higher-sensitivity projects.
• Competitive compensation.
• Opportunity to collaborate with leading researchers.
• Referral bonus of up to $250 for each successful referral.
Cloudera
NextLink Group
Thrive
Get handpicked remote jobs straight to your inbox weekly.