
AI Safety Expert, English β Dutch
Posted 5 hours ago

Posted 5 hours ago
This is a fully remote position, open to applicants in United States.
β’ Engage in red team assessments of conversational AI models and agents.
β’ Test for jailbreaks, prompt injections, misuse scenarios, bias exploitation, and multi-turn manipulations.
β’ Create human data by annotating failures, identifying vulnerabilities, and highlighting systemic risks.
β’ Utilize taxonomies, benchmarks, and playbooks to ensure consistent testing methodologies.
β’ Generate reproducible reports, datasets, and actionable attack scenarios.
β’ Investigate AI outputs related to sensitive subjects such as bias, misinformation, and harmful behaviors.
β’ Discover vulnerabilities that automated testing may overlook.
β’ Broaden evaluation coverage and minimize surprises in production.
β’ Enhance customer AI systems through adversarial testing.
β’ Must possess native fluency in both English and Dutch.
β’ Previous experience in red teaming within AI adversarial contexts, cybersecurity, or socio-technical probing is essential.
β’ Capability to conduct adversarial probing of AI systems, including jailbreaks, prompt injections, misuse cases, bias exploitation, and multi-turn manipulation.
β’ Proficient in annotating failures, classifying vulnerabilities, and identifying systemic risks.
β’ Experience in adhering to taxonomies, benchmarks, and playbooks.
β’ Ability to create reproducible reports, datasets, and attack scenarios.
β’ Skill in articulating risks clearly to both technical and non-technical stakeholders.
β’ Must be adaptable across various projects and clients.
β’ H1-B and STEM OPT candidates are not eligible for support.
β’ This position is for independent contractor engagement.
β’ Fully remote position.
β’ Flexible scheduling tailored to your availability.
β’ Weekly payments processed through Stripe or Wise.
β’ Project durations may vary based on needs and performance metrics.
β’ Participation in higher-sensitivity projects is optional.
β’ Clear guidelines and wellness resources are provided for handling sensitive content.
β’ Up to $250 referral bonus for each successful referral.
β’ Competitive compensation package.
β’ Opportunity to collaborate with leading researchers in the field.
β’ Reasonable accommodations available upon request.
XenoPatch GmbH
Mercor
Mercor
Mercor
Get handpicked remote jobs straight to your inbox weekly.