
AI Safety Expert β English, Dutch
Posted Sep 3

Posted Sep 3
This is a fully remote position, open to applicants in United States.
β’ Engage in red team assessments of conversational AI models and agents, utilizing techniques such as jailbreaks, prompt injections, misuse scenarios, bias exploitation, and multi-turn manipulation.
β’ Document failures, categorize vulnerabilities, and identify systemic risks.
β’ Adhere to established taxonomies, benchmarks, and playbooks to ensure consistent testing practices.
β’ Generate reproducible reports, datasets, and attack scenarios.
β’ Investigate sensitive topics such as bias, misinformation, and harmful behaviors.
β’ Provide deliverables that enhance evaluation coverage and fortify customer AI systems.
β’ Collaborate with top researchers on initiatives aimed at training and improving AI systems.
β’ Required native fluency in both English and Dutch.
β’ Previous red teaming experience in AI adversarial contexts, cybersecurity, or socio-technical investigations.
β’ Capability to test AI systems adversarially and push them to their limits.
β’ Proficient in utilizing frameworks, taxonomies, benchmarks, and playbooks for systematic testing.
β’ Ability to clearly communicate risks to both technical and non-technical audiences.
β’ Flexibility to adapt to various projects and client needs.
β’ Must maintain independent contractor status.
β’ H1-B and STEM OPT candidates are not eligible for this position.
β’ Fully remote work environment.
β’ Flexible scheduling to suit your needs.
β’ Weekly payments processed via Stripe or Wise.
β’ Competitive compensation.
β’ Optional participation in higher-sensitivity projects.
β’ Comprehensive guidelines and wellness resources for projects involving sensitive content.
β’ Referral bonus of up to $250 for each successful referral.
The College Board
Mercor
Mercor
Get handpicked remote jobs straight to your inbox weekly.